Spectral-Driven Vibrato — User Guide

Global spectral analysis controls a causal fractional-delay vibrato: spectral flatness sets the vibrato depth, while normalized spectral spread sets its rate.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 0.4 (2026) License: MIT License Repo: https://github.com/ShaiCohen-ops/Praat-plugin_AudioTools
Contents:

What this does

Spectral-Driven Vibrato derives one vibrato setting from the overall spectrum of the selected Sound. A mono analysis copy is windowed and transformed to a spectrum over the requested frequency band. Two global spectral descriptors are then measured:

Those two derived values remain fixed for the duration of the Sound and drive a sinusoidally moving causal delay line. The audio itself is processed from the original Sound, so the spectral analysis acts as a controller for the vibrato rather than replacing or resynthesizing the source.

In practical terms: a flatter spectrum produces more depth within the selected mapping range, while a wider spectral distribution produces a faster vibrato. The source spectrum therefore determines the character of one global vibrato applied to the recording.

Processing pipeline

1. Validate the selected Sound and requested analysis band.
2. Build a mono analysis signal, with a cancellation fallback for multichannel sources.
3. Apply a Hann window to the disposable analysis copy.
4. Compute one exact-length FFT spectrum for the complete Sound.
5. Measure flatness, centroid, and spectral spread inside the requested band.
6. Map flatness to vibrato depth and normalized spread to vibrato rate.
7. Convert the requested semitone depth into a sinusoidal delay excursion.
8. Apply causal fractional-delay vibrato independently to every original channel.
9. Apply attenuation-only peak safety when required.
10. Optionally draw the AudioTools visualization and play the result.

The spectral measurements are file-level values. The script does not divide the recording into spectral frames for the audio mapping; instead, one global analysis determines one depth and one rate for the complete result.

Global spectral analysis

Analysis channel selection

A mono Sound is analyzed directly. A multichannel Sound is first converted to mono. The script also measures the absolute peak of each original channel. If the mono fold-down peak is less than 10% of the strongest channel peak, the strongest individual channel is used instead. This prevents strong opposite-polarity channel content from largely cancelling the signal that drives the analysis.

Window and FFT

The complete disposable analysis copy receives a Hann window:

w[n] = 0.5 − 0.5 cos(2πn / (N−1))

Praat then creates an exact-length spectrum. Only FFT bins between Min frequency and the resolved Max frequency are included in the descriptor calculations.

Spectral flatness

For every included bin, power is calculated from the real and imaginary Spectrum cells:

P[k] = Re[k]² + Im[k]² flatness = exp( mean( ln(P[k]) ) ) ----------------------- mean(P[k])

Power is floored at a very small positive value before the logarithm. The final flatness value is constrained to 0…1. In this processor, flatness is the control variable for vibrato depth.

Spectral centroid and spread

centroid = Σ(f[k] P[k]) / ΣP[k] spread = sqrt( Σ(f[k]² P[k]) / ΣP[k] − centroid² )

The centroid is calculated and reported as descriptive information. The vibrato rate is driven by spectral spread.

Normalized spread

The raw spread in Hz is normalized relative to the width of the selected analysis band:

analysisWidth = Max_frequency − Min_frequency spreadNorm = clamp( sqrt(12) × spread / analysisWidth, 0, 1 )

Spread Response then scales this normalized value before the rate mapping.

If the analysis signal is effectively silent, flatness, centroid, spread, and normalized spread are set to zero. The mapping therefore falls back to the selected Base Depth and Base Rate values.

Spectral-to-vibrato mapping

Depth from spectral flatness

depth_st = clamp( Base_depth + flatness × Max_depth_add, 0, 2 semitones )

Base Depth is the minimum mapped depth. Max Depth Add determines how much additional depth a flatness value of 1 can contribute.

Rate from spectral spread

spreadDrive = clamp( spreadNorm × Spread_response, 0, 1 ) rate_Hz = clamp( Base_rate_Hz + spreadDrive × Max_rate_add_Hz, 0, 50 Hz )

Spread Response changes how quickly normalized spread reaches the top of the rate mapping. Values above 1 make the rate mapping reach full drive earlier; values below 1 reduce its response.

Control summary

Flatness → Depth
Normalized spread → Rate
Centroid → reported measurement only

Causal fractional-delay vibrato

The derived depth is expressed in semitones, while the actual DSP is a moving delay line. For a sinusoidal delay trajectory, the local pitch ratio is approximately determined by the derivative of the delay. v0.4 chooses the delay excursion so that the positive pitch peak corresponds exactly to the requested mapped depth under this model.

Semitone depth to delay excursion

positive pitch-ratio excursion = 2^(depth_st / 12) − 1 delayExcursion = pitch-ratio excursion / (2π × rate_Hz)

Delay trajectory

D(t) = effectiveBaseDelay − delayExcursion × cos(2π × rate_Hz × local_time)

The source is read at t − D(t) using Praat's time-based Sound interpolation, producing fractional-delay modulation rather than integer-sample stepping.

Causality and short-file protection

The script limits the maximum usable delay to approximately 45% of the Sound duration, with a minimum allowance tied to the sample interval. If the requested modulation would require too much delay, the delay excursion is reduced and the effective depth is recalculated. The base delay can also be raised automatically to keep the complete modulation trajectory safely causal.

Initial delay-line fill

At the beginning of the Sound, the wet contribution is faded in over the maximum delay time. This lets the delay line fill from valid source history instead of abruptly starting at full wet level.

Dry/Wet

After the initial fill ramp, the normal blend is:

output = (1 − Wet) × original + Wet × delayed_original

0% Wet is an exact bypass. The same bypass path is also used when the derived depth or rate is zero.

Output safety

After active processing, the result peak is measured. If it exceeds Safety Peak, the complete Sound is scaled down to that ceiling. Material already below the ceiling is left unchanged.

Safety Peak is attenuation-only. It is not target normalization and is skipped on the exact bypass path.

Presets

The named presets change the depth/rate mapping controls. The analysis frequency range, base delay, Dry/Wet, Safety Peak, visualization, and playback remain under the user's direct control.

Preset Base depth Max depth add Base rate Max rate add Spread response
Custom0.05 st0.15 st4.0 Hz3.0 Hz1.00
Subtle Natural0.03 st0.10 st5.0 Hz2.0 Hz1.00
Moderate Expressive0.05 st0.15 st4.5 Hz3.0 Hz1.00
Strong Character0.08 st0.25 st4.0 Hz4.0 Hz1.00
Fast Flutter0.04 st0.10 st6.0 Hz4.0 Hz1.25
Slow Sweep0.10 st0.20 st2.5 Hz2.0 Hz0.85

Parameters

ParameterDefaultWhat it controlsRuntime handling
Min frequency Hz80Lower edge of the global spectral-analysis band.Minimum 0 Hz.
Max frequency Hz5000Upper edge of the analysis band.Limited to 0.98 × Nyquist.
Spread response1.0Scales normalized spread before rate mapping.Clamped to 0…4.
Base depth0.05 stMinimum depth before flatness contribution.Clamped to 0…2 semitones.
Max depth add0.15 stMaximum flatness-controlled addition.Clamped to 0…2 semitones.
Base rate Hz4.0Minimum vibrato rate before spread contribution.Clamped to 0…50 Hz.
Max rate add Hz3.0Maximum spread-controlled rate addition.Clamped to 0…50 Hz.
Base delay ms5.0Requested center delay of the vibrato trajectory.Minimum 0; can be increased automatically to maintain causality.
Dry/Wet percent100Blend between direct and fractionally delayed source.Clamped to 0…100%; 0 gives exact bypass.
Safety peak0.99Maximum peak after active processing.Clamped to 0…1; 0 disables scaling; attenuation only.
Draw visualizationyesCreates the AudioTools figure.Does not change the audio.
Play resultyesPlays the completed Sound.Does not change the audio.

Channels & output

PropertyBehavior
Analysis signalMono input directly; otherwise mono fold-down with strongest-channel fallback if the fold-down is strongly cancelled.
Audio processingPerformed on every channel of the original Sound using the same derived depth, rate, and delay trajectory.
Channel countPreserved.
Sample ratePreserved.
Sample countPreserved.
DurationPreserved.
Start timePreserved.
LFO phaseUses local time from the Sound start, so moving an otherwise identical Sound on the Praat timeline does not change the modulation trajectory.
RandomnessNone. The same source and settings produce the same analysis and output.

Visualization

When Draw visualization is enabled, v0.4 draws the AudioTools house layout:

Input

The original Sound waveform.

Output

The final processed waveform after optional Safety Peak attenuation.

Analysis spectrum

The global spectrum inside the requested analysis band. Up to 300 representative bins are retained for drawing; the descriptor calculations use every valid bin in the band.

Pitch modulation

The approximate pitch-ratio trajectory converted to semitones over the first up to 1 second. The zero line marks unshifted pitch.

Summary strip

The summary reports the measured flatness and normalized spread, the resulting depth and rate, analysis frequency range, Wet percentage, Safety setting, duration, and channel count.

The spectrum panel shows the single global analysis that controls the effect. The pitch panel shows the vibrato trajectory produced from the resulting global depth and rate.