Spectral-Driven Intensity Modulation — User Guide

Time-varying spectral analysis drives an attenuation-only tremolo: spectral flatness controls depth, normalized spectral spread controls rate, and tonal compact-spectrum regions can be protected by reducing both.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 0.4 (2026) License: MIT License Repo: https://github.com/ShaiCohen-ops/Praat-plugin_AudioTools
Contents:

What this does

Spectral-Driven Intensity Modulation listens to how the spectrum changes through the selected Sound and turns those changes into a continuously evolving tremolo. Short analysis windows are distributed across the recording. Each window measures two scale-independent spectral descriptors:

The measured values are interpolated between analysis windows, so depth and rate evolve continuously rather than jumping from one window to the next. When a region is both sufficiently tonal and spectrally compact, optional protection factors reduce the derived depth and rate before the tremolo is generated.

The effect is attenuation-only. Its gain trajectory runs from 0 dB down to a negative depth value and never boosts as part of the tremolo itself. Dry/Wet then blends that attenuated version with the original Sound.

Processing pipeline

1. Validate the Sound, spectral band, modulation ranges, and advanced controls.
2. Choose a mono analysis signal, with strongest-channel fallback if a multichannel fold-down is strongly cancelled.
3. Place short Hamming-windowed analysis regions across the recording.
4. Compute an exact-length Spectrum for every analysis window.
5. Measure spectral flatness, centroid, and normalized spectral spread inside the requested band.
6. Interpolate flatness and spread through time.
7. Map flatness to depth and spread to modulation rate, with optional tonal/compact-spectrum protection.
8. Integrate the changing rate into a continuous modulation phase.
9. Write a relative-dB IntensityTier and multiply it with the original multichannel Sound.
10. Apply Dry/Wet mixing and an attenuation-only Safety Peak ceiling.

The spectral features control the effect; they do not replace the source audio. The final Sound retains the original sample grid and channel structure.

Spectral analysis

Analysis source

Mono input is copied directly for analysis. Multichannel input is first converted to mono. The script also measures every original channel independently. If the mono fold-down peak is less than 10% of the strongest channel peak, the fold-down is considered vulnerable to phase cancellation and the strongest individual channel is used for analysis instead.

The analysis copy is shifted to a local 0-second start when the original Sound has a non-zero xmin. This simplifies analysis timing without altering the original Sound or its eventual output time domain.

Window placement

By default, 8 analysis windows of 200 ms are distributed evenly across the usable duration. Their centers run from half a window after the beginning to half a window before the end.

If the Sound is shorter than the requested window duration, the complete Sound is reused for every spectral measurement while the control timestamps are still distributed over the valid output duration.

Frequency band

The default analysis band is 80–5000 Hz. The upper edge is clamped to 0.49 × sample rate. The script requires at least two valid FFT bins inside the resolved band.

Spectral flatness

power[k] = Re[k]² + Im[k]² meanPower = mean(power[k]) relativeFloor = max(1e-300, meanPower × 1e-12) flatness = exp( mean( ln(max(relativeFloor, power[k])) ) ) ------------------------------------------------ meanPower

The floor is relative to the mean spectral power, so multiplying the source by a constant level does not materially change the flatness estimate. The final value is clamped to 0…1.

Spectral centroid and normalized spread

centroid = Σ(f[k] × power[k]) / Σpower[k] spread = sqrt( Σ(f[k]² × power[k]) / Σpower[k] − centroid² ) spreadNorm = clamp( sqrt(12) × spread / analysisBandWidth, 0, 1 )

The centroid is measured but is not used as a modulation controller. The rate mapping uses normalized spectral spread. The factor sqrt(12) makes a uniform spectral distribution across the complete analysis band approach a normalized spread of 1.

If the analysis signal is effectively silent, flatness, spread, and centroid are set to zero for every analysis point. The tonal/compact-spectrum protection test can then reduce the base depth and base rate according to its reduction factors.

Spectral mapping

Flatness → attenuation depth

depth_dB(t) = Base_depth_dB + flatness(t) × [Max_depth_dB − Base_depth_dB]

Before tonal protection, flatness 0 maps to Base Depth and flatness 1 maps to Max Depth.

Spread → modulation rate

rate_Hz(t) = Base_mod_speed_Hz + spreadNorm(t) × [Max_mod_speed_Hz − Base_mod_speed_Hz]

Before tonal protection, normalized spread 0 maps to Base Rate and spread 1 maps to Max Rate.

Tonal / compact-spectrum protection

The protection stage activates only when both conditions are true:

flatness(t) < Tonal_flatness_threshold AND spreadNorm(t) < Smooth_spread_threshold

When active:

depth_dB(t) = depth_dB(t) × Tonal_depth_reduction rate_Hz(t) = rate_Hz(t) × Tonal_speed_reduction

Because the reduction is applied after the normal mapping, protected regions can fall below the nominal Base Depth and Base Rate. With the default factors, the depth is multiplied by 0.3 and the rate by 0.7.

Gain trajectory

Adaptive control resolution

The Advanced Time step is an upper limit, not always the final spacing. The script also guarantees at least 32 control points per cycle of the fastest requested modulation rate:

controlStep = min( Time_step, 1 / [32 × Max_mod_speed_Hz] )

With the default 5 Hz maximum rate and a 10 ms requested step, this resolves to 6.25 ms.

Interpolation of spectral controls

At every control point, flatness and normalized spread are linearly interpolated between the two surrounding spectral-analysis timestamps. Before the first analysis center, the first measurement is held; after the last center, the final measurement is held.

Continuous phase

Changing rate is converted into one continuous phase accumulator:

phase[i] = phase[i−1] + 2π × rate_Hz[i] × Δt

This avoids restarting the oscillator whenever the spectrally derived rate changes.

Attenuation-only tremolo law

gain_dB(t) = −0.5 × depth_dB(t) × [1 − cos(phase(t))]

The range is exactly 0 dB to −depth_dB. Phase starts at zero, so the processed path begins at 0 dB rather than with an artificial level discontinuity.

Relative-dB IntensityTier

The gain values are written directly to an IntensityTier in the original Sound's absolute time domain. Praat interprets IntensityTier values as relative dB multipliers. The script calls Multiply: "no", explicitly disabling Praat's optional post-multiply scaling to 0.9.

Dry/Wet

output = Wet × processed + Dry × original

Since the processed path is itself attenuation-only, the Dry/Wet blend also cannot exceed the instantaneous original amplitude through the tremolo mapping alone. 0% Wet is an exact copied bypass.

Safety Peak

After active processing, the script measures the peak. If it exceeds Safety Peak, the complete Sound is scaled down to that ceiling. Signals already below the threshold are not raised. Safety Peak is skipped on the explicit 0% Wet bypass.

Presets

Preset Depth range Rate range Tonal protection changes
Custom20–50 dB1–5 HzDefault advanced values
Subtle Texture12–30 dB0.8–3 HzDefault advanced values
Moderate Dynamics18–42 dB1–4 HzDefault advanced values
Strong Spectral Response22–52 dB1.5–6 HzDefault advanced values
Voice Protection Mode16–36 dBdefault 1–5 HzFlatness threshold .40; spread threshold .18; depth ×.20; speed ×.50
Maximum Effect26–60 dB2–8 HzDepth reduction becomes ×.60; other advanced protection defaults remain
Presets override only the values explicitly assigned by the script. In particular, Voice Protection changes several Advanced classifier/protection values, while Maximum Effect changes the tonal depth-reduction factor but leaves the tonal thresholds and speed-reduction factor at their current values.

Parameters

Main form

ParameterDefaultMeaningRuntime handling
PresetCustomSelects one of five mapped behaviors or the manual values below.Named presets overwrite their assigned mapping values.
Base depth dB20Depth mapped from flatness = 0 before protection.Clamped to 0…80 dB.
Max depth dB50Depth mapped from flatness = 1 before protection.Clamped to Base Depth…80 dB.
Base mod speed Hz1.0Rate mapped from spread = 0 before protection.Clamped to 0…50 Hz.
Max mod speed Hz5.0Rate mapped from spread = 1 before protection.Clamped to Base Rate…50 Hz.
Dry/Wet percent100Global blend of original and attenuated path.Clamped to 0…100%; 0 gives exact bypass.
Advanced settingsoffOpens the analysis, tonal-protection, resolution, and safety controls.—
Draw visualizationyesCreates the AudioTools figure.Does not change audio.
Play resultyesPlays the completed result.Does not change audio.

Advanced settings

ParameterDefaultMeaningRuntime handling
Num analysis points8Number of short spectral measurements across the Sound.Clamped to 2…64.
Window size seconds0.2Length of each Hamming-windowed spectral analysis region.Clamped to .005 s…Sound duration.
Min frequency Hz80Lower analysis-band edge.Clamped to 0…0.49 × sample rate.
Max frequency Hz5000Upper analysis-band edge.Clamped to 0…0.49 × sample rate; must remain above Min.
Tonal flatness threshold.30Upper flatness limit for tonal protection.Clamped to 0…1.
Smooth spread threshold.12Upper normalized-spread limit for tonal protection.Clamped to 0…1.
Tonal depth reduction.30Multiplier applied to depth when both protection tests pass.Clamped to 0…1.
Tonal speed reduction.70Multiplier applied to rate when both protection tests pass.Clamped to 0…1.
Time step.01 sMaximum spacing requested for the IntensityTier control grid.Clamped to .0005….05 s, then possibly reduced to satisfy 32 points/cycle.
Safety peak.99Maximum peak allowed after active processing.Clamped to 0…1; 0 disables safety; attenuation only.

Channels & output

PropertyBehavior
AnalysisMono input directly, otherwise mono fold-down with strongest-channel fallback when cancellation is severe.
Gain controlOne shared relative-dB IntensityTier controls every source channel identically.
Channel countPreserved, including arbitrary multichannel Sounds.
Sample ratePreserved.
DurationPreserved.
Start timePreserved; the control tier is created directly over the original absolute time domain.
RandomnessNone. Analysis and modulation are deterministic for the same source and settings.
0% WetExact copied bypass.
Maximum active gain0 dB before Safety Peak; the tremolo itself never boosts.

Visualization

When Draw visualization is enabled, v0.4 produces the standardized AudioTools picture:

Input

The original waveform.

Output

The final processed waveform after optional safety attenuation.

Spectral drives

Interpolated spectral flatness and normalized spread over time. The dotted horizontal line marks the tonal-flatness threshold.

Gain modulation

The actual relative-dB control trajectory, from 0 dB downward.

Visualization decimation

The full IntensityTier can contain many more points, especially at high modulation rates. The figure stores at most 500 representative control points for drawing. This does not reduce the resolution of the audio modulation itself.

Summary strip

The summary reports average interpolated flatness, average normalized spread, average derived depth, average derived rate, Wet percentage, Safety setting, duration, sample rate, and channel count.

The spectral-drive panel shows the controller variables; the gain-modulation panel shows the actual attenuation command generated from those variables. The latter is the quantity ultimately written to the IntensityTier.