Partial Editing Resynthesis — Sinusoidal Texture Resynthesis — User Guide
Frame-based additive resynthesis that selects strong spectral partials, edits their frequency and amplitude, and reconstructs a new mono texture.
What this does
Partial Editing Resynthesis analyzes short overlapping frames, finds a limited number of strong spectral peaks, edits their frequencies and amplitudes, and rebuilds the sound as a sum of sine waves. The script's form is titled Sinusoidal Texture Resynthesis because the result is intentionally a new additive texture rather than a transparent reconstruction of the source.
For each frame, the processor searches between Min_frequency and Max_frequency, repeatedly takes the strongest remaining spectral peak, suppresses a ±40 Hz region around that peak, and continues until Max_partials_per_frame has been reached or no sufficiently strong peak remains.
Sinusoidal analysis and resynthesis
A sinusoidal model represents sound as a collection of time-varying sine components. Here the analysis is deliberately simpler and more textural than a full sinusoidal-tracking model:
- Peaks are detected independently in each analysis frame.
- Peak frequencies are kept on FFT-bin centres before user editing.
- There is no trajectory tracking that follows one partial from frame to frame.
- The source phase is not copied into the new oscillators; resynthesis uses sine generators evaluated on absolute time.
- There is no separate stochastic/noise residual model.
Overlapping Hann-windowed grains are summed and then divided by the exact accumulated window envelope. This gives stable overlap-add reconstruction even when the user chooses a window/hop ratio that is not a textbook constant-overlap-add setting.
Quick start
- Select exactly one Sound. Mono and multichannel sources are accepted, but the output is always mono.
- Run
Partial_Editing_Resynthesis.praat. - Start with Clean Texture Resynth or Dense Partials.
- Use Max_partials_per_frame to move between sparse and rich reconstructions.
- Use Transpose_semitones for musical transposition and Additional_frequency_scale for an additional multiplicative spectral scale.
- Choose stable jitter for continuity or frame-random jitter for the older grainier texture.
Presets
| Preset | Key overrides |
|---|---|
| Custom | No override. |
| Clean Texture Resynth | Freq jitter 0.5 Hz; amp jitter 0.02; 20 partials. |
| Diffuse Texture | 8 Hz; 0.3; 15 partials. |
| Sparse Partials | 5 partials; 2 Hz; 0.1. |
| Dense Partials | 30 partials; 1 Hz; 0.05. |
| Pitch Up Octave | +12 st; 1 Hz; 0.05. |
| Pitch Down Octave | −12 st; 1 Hz; 0.05. |
| Spectral Scale Up | Additional frequency scale 1.5; 2 Hz; 0.1. |
| Spectral Scale Down | Additional frequency scale 0.7; 2 Hz; 0.1. |
| Glassy Shimmer | 15 Hz; 0.4; 20 partials; max frequency 12 kHz (clipped to Nyquist if necessary). |
| Robotic | No frequency or amplitude jitter; 12 partials. |
| Whisper Ghost | 4 partials; 10 Hz; 0.5; max frequency 6 kHz. |
Controls
| Control | Default | Meaning |
|---|---|---|
| Window_length | 0.060 s | Duration of each Hann analysis/resynthesis frame. |
| Hop_size | 0.015 s | Time between frame starts. The final frame is shifted when necessary so the source end is covered. |
| Min_frequency / Max_frequency | 60 / 8000 Hz | Peak-search band. Max is clipped to source Nyquist. |
| Max_partials_per_frame | 15 | Maximum number of strongest peaks resynthesized in each frame. |
| Freq_jitter | 3 Hz | Maximum detuning amount around each detected FFT-bin frequency; negative values are clamped to 0. |
| Amp_jitter | 0.1 | Relative amplitude variation; internally clamped to 0…1. |
| Jitter_behaviour | Stable per-frequency | Stable mode deterministically maps each FFT bin to a fixed detune/gain offset. Frame-random mode generates new random offsets in each frame. |
| Transpose_semitones | 0 | Frequency multiplier 2^(st/12). |
| Additional_frequency_scale | 1.0 | Additional positive multiplier applied to every resynthesized partial frequency. |
The final partial frequency is approximately (detected_frequency + jitter) × 2^(transpose/12) × additional_scale. Components below 20 Hz or at/above working Nyquist are omitted from the synthesized grain.
Processing pipeline
- For multichannel input, measure each source channel and choose the highest-RMS channel as the analysis driver. This avoids cancellation that could occur if anti-phase channels were averaged.
- Create a mono output buffer and a separate overlap-add window-sum buffer at the original sample rate and duration.
- For every frame: extract a Hann-windowed segment, compute its Spectrum and Ltas, find strong peaks one by one, and suppress ±40 Hz around each chosen peak before searching again.
- Apply the selected frequency/amplitude jitter, transposition and additional frequency scaling.
- Synthesize all accepted partials for that frame as sine waves under an analytic Hann window.
- Overlap-add every grain into the output.
- Divide by the exact accumulated Hann-window envelope.
- If the result is non-silent, apply
Scale intensity: 70.
Channels, duration, randomness and level
- Output: always mono.
- Multichannel analysis: highest-RMS source channel only; the other source channels are not mixed into the synthesis.
- Sample rate: source sample rate.
- Duration: source duration.
- Stable jitter: deterministic for a given FFT-bin identity and parameter set; no random seed is required.
- Frame-random jitter: uses unseeded random values, so repeated runs can differ.
- Final level: non-silent output is set to 70 dB SPL in Praat's intensity convention. This is not 70 dBFS and is not peak normalization.
Output name: <source>_resynth_<preset>.
Visualization
The figure shows original and resynthesized waveforms on one shared amplitude scale, original/output spectrograms, and a summary of partial count, jitter, transposition and frequency scale.
Historical, technological and compositional context
Sinusoidal analysis/resynthesis became a major digital sound-modeling technique in the 1980s. McAulay and Quatieri's 1986 sinusoidal representation estimated short-time sinusoidal frequencies, amplitudes and phases from spectral peaks and tracked those components through time. In computer music, Serra and Smith's 1990 Spectral Modeling Synthesis extended spectral modeling by separating a sound into deterministic sinusoidal trajectories and a stochastic residual.
This Praat tool deliberately takes a more reduced, compositional approach. It keeps only a limited set of strong peaks in each frame, does not track partial trajectories across frames, and does not resynthesize a noise residual. Those omissions are not merely technical limitations: they are what make sparse, robotic, diffuse and ghost-like textures possible. The source is treated as material for a new additive reconstruction rather than as something that must be reproduced transparently.
Further reading
McAulay, R. J., & Quatieri, T. F. (1986). “Speech Analysis/Synthesis Based on a Sinusoidal Representation.” IEEE Transactions on Acoustics, Speech, and Signal Processing, 34(4), 744–754.
Serra, X., & Smith, J. (1990). “Spectral Modeling Synthesis: A Sound Analysis/Synthesis Based on a Deterministic plus Stochastic Decomposition.” Computer Music Journal, 14(4), 12–24.