Spectral Freeze Synthesis — User Guide

Evolving spectral peak-hold freeze: dominant peaks are retained, decayed, replaced by stronger peaks, optionally glissandoed, and additively resynthesized.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 0.8.1 (2026) License: MIT License Repo: https://github.com/ShaiCohen-ops/Praat-plugin_AudioTools
Contents:

What this does

Spectral Freeze Synthesis is an evolving peak-hold resynthesizer. It repeatedly analyzes short frames of the source, extracts the strongest spectral peaks, and maintains a bank of held frequencies and amplitudes. Each held amplitude decays over time; a newly measured peak replaces a held slot only when it becomes stronger than that slot's decayed value. The current bank is then resynthesized additively as overlapping Hann-windowed sine grains.

The result can behave like a static drone, a slowly disintegrating spectrum, or a continuously rising/falling spectral cloud. Unlike a conventional FFT freeze that simply repeats one complete analysis frame, this script keeps an evolving ranked bank of dominant peaks.

What is spectral freeze? In digital audio, “spectral freeze” usually means holding spectral information so that a short moment becomes sustained. Many phase-vocoder implementations hold an STFT/FFT frame or its spectral parameters. This script uses a different mechanism: it selects only a limited number of dominant peaks, lets their amplitudes decay, allows stronger new peaks to replace them, and resynthesizes the bank as sine waves.

Peak-hold mechanism

For every analysis frame, the processor searches for the strongest peak below max_frequency_hz, suppresses a neighbourhood around it, then searches again until top_partials peaks have been considered. The suppression width is derived from peak_separation_hz, so nearby peaks are less likely to occupy multiple slots.

Per-frame retention: d_frame = decay_factor ^ frame_step_seconds Per-frame drift: gliss_ratio = 2 ^ (glissando_oct_sec * frame_step_seconds) For each held slot k: amplitude[k] *= d_frame frequency[k] *= gliss_ratio clamp frequency to 20 Hz … max_frequency_hz If current ranked peak k is stronger than the decayed held amplitude: replace that slot with the new peak.

decay_factor is therefore a per-second retention ratio, not a per-frame multiplier. A value of 1.0 is an infinite hold; smaller values decay faster. glissando_oct_sec is measured in octaves per second: positive values drift upward, negative values downward.

Phase character: the script intentionally preserves its legacy grain-local phase-reset behaviour. Each short additive grain is regenerated from sine oscillators rather than carrying continuous source phase from one frame to the next. This is part of the processor's characteristic texture.

Quick start

  1. Select exactly one Sound.
  2. Run Spectral_Freeze_Synthesis.praat.
  3. Start with Classic Freeze for a stable held bank, Gentle Decay for a slowly changing texture, or Rising Shimmer for upward drift.
  4. Use top_partials to control spectral density and peak_separation_hz to control how closely selected peaks may cluster.
  5. For faster processing, lower resample_to_hz; remember that this also lowers the wet path's Nyquist limit.
  6. Use tail_duration_sec to allow the held bank to continue after the source ends.

Presets

Presets override only the values listed below. Other form values remain as entered by the user.

PresetOverridesCharacter
CustomNoneUser settings.
Classic Freezedecay 1.0; gliss 0; 10 partialsIndefinite held-bank sustain until replaced by stronger peaks.
Gentle Decay0.5; 0; 10Slowly fading held spectrum.
Rising Shimmer0.3; +0.15 oct/s; 12Held components drift upward.
Falling Shimmer0.3; −0.15 oct/s; 12Held components drift downward.
Ghostly Fade0.15; +0.02; 20; tail 4 sDense but quickly fading cloud.
Metallic Drone0.95; 0; 6; max 3000 HzSparse, persistent low/mid spectral bank.
Cosmic Drift0.4; +0.5; 15; tail 5 sLarge upward glissando.
Frozen Choir0.8; 0; 16; stereo delay 20 ms; max 4000 HzDense freeze with wider wet-channel phase offset.
Disintegrating0.05; −0.05; 25; tail 1 sMany components that vanish rapidly.
Ascending to Heaven0.6; +0.8; 10; tail 4 s; max 5000 HzStrong upward spectral drift.
Descending to Hell0.7; −0.6; 8; tail 4 s; max 2500 HzStrong downward drift concentrated lower in the spectrum.
Spectral Dust0.02; +0.1; 3; tail 0.5 sVery sparse, short-lived particles.

Controls

ControlDefaultMeaning
frame_step_ms20 msDistance between analysis/resynthesis frames.
analysis_window_ms35 msHann analysis window and additive grain duration.
max_frequency_hz6000Upper peak-search and held-frequency limit; capped to the work-rate Nyquist frequency.
top_partials10Number of ranked held slots; values below 1 are forced to 1.
peak_separation_hz50 HzApproximate exclusion width around each selected peak.
resample_to_hz12000 HzWorking sample rate for wet analysis and synthesis. Lower values speed processing but remove frequencies above work Nyquist.
restore_original_sample_rateOnResamples the synthesized wet output back to the source sample rate. It does not restore high frequencies discarded by low-rate analysis.
decay_factor0.2Per-second amplitude retention ratio, capped at 1.
glissando_oct_sec0Exponential frequency drift in octaves per second.
wet_dry_percent100%Effect blend, clamped to 0…100.
tail_duration_sec2 sExtra analysis/synthesis time after the source, allowing the held bank to continue.
create_stereo_outputOnCreates a two-channel wet synthesis. When off, the wet synthesis is mono.
stereo_delay_ms8 msApplied inside the right sine formula as a frequency-dependent phase offset equivalent to a delay; it is not implemented as literal leading silence.
target_peak_db−1 dBWhen Wet > 0, non-silent final output is target-normalized to 10^(dB/20).
draw_visualizationOnDraws measured held-bank and source/wet/final QC views.
play_afterOnPlays the final result after processing.

Processing pipeline

  1. Resample a multichannel work copy to resample_to_hz.
  2. Convert a separate analysis copy to mono. Peak detection is global; the original dry channel structure is kept separately.
  3. Append the requested tail to the mono analysis signal.
  4. For each frame, compute a magnitude spectrum, find ranked peaks with spectral suppression, decay/drift the held bank, and replace slots only when the new peak is stronger.
  5. Resynthesize all active held components as Hann-windowed additive grains and overlap-add them into the wet buffer.
  6. Optionally resample the wet output back to the original sample rate.
  7. Mix dry and wet. Beyond the source duration, the dry signal evaluates to zero, so the tail is wet-only.
  8. If Wet > 0 and the result is non-silent, target-normalize to target_peak_db.

Channels, bypass, duration and level

Visualization

The visualization is measurement-driven. It records the actual held-bank frequency and amplitude state after every analysis frame, so the trajectory panel is a QC view of the states that drove synthesis rather than an illustrative sketch. Additional panels compare representative source, pure wet and final output measurements.

Technological and compositional context

Spectral freezing belongs to the broader family of frequency-domain sound transformations that became central to computer music through STFT and phase-vocoder techniques. Mark Dolson's 1986 phase-vocoder tutorial framed spectral analysis/resynthesis as a compositional tool capable of manipulating individual spectral components and time/pitch structure.

This processor is not a conventional phase-vocoder frame freeze. Instead of holding a complete magnitude/phase frame, it reduces each frame to a ranked set of strong peaks, carries those peaks forward through decay and glissando, and rebuilds them additively. Compositionally, this turns a momentary spectrum into an evolving harmonic/inharmonic “memory” whose contents may persist, drift, be replaced, or dissolve.

Further reading

Dolson, M. (1986). “The Phase Vocoder: A Tutorial.” Computer Music Journal, 10(4), 14–27. DOI: 10.2307/3680093.