Partial Editing Resynthesis — Sinusoidal Texture Resynthesis — User Guide

Frame-based additive resynthesis that selects strong spectral partials, edits their frequency and amplitude, and reconstructs a new mono texture.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 0.7 (2026) License: MIT License Repo: https://github.com/ShaiCohen-ops/Praat-plugin_AudioTools
Contents:

What this does

Partial Editing Resynthesis analyzes short overlapping frames, finds a limited number of strong spectral peaks, edits their frequencies and amplitudes, and rebuilds the sound as a sum of sine waves. The script's form is titled Sinusoidal Texture Resynthesis because the result is intentionally a new additive texture rather than a transparent reconstruction of the source.

For each frame, the processor searches between Min_frequency and Max_frequency, repeatedly takes the strongest remaining spectral peak, suppresses a ±40 Hz region around that peak, and continues until Max_partials_per_frame has been reached or no sufficiently strong peak remains.

What is a partial? A partial is one sinusoidal component of a complex sound. Harmonic sounds often contain partials near integer multiples of a fundamental frequency, but this processor does not assume harmonicity: it simply selects strong short-time spectral peaks.

Sinusoidal analysis and resynthesis

A sinusoidal model represents sound as a collection of time-varying sine components. Here the analysis is deliberately simpler and more textural than a full sinusoidal-tracking model:

Overlapping Hann-windowed grains are summed and then divided by the exact accumulated window envelope. This gives stable overlap-add reconstruction even when the user chooses a window/hop ratio that is not a textbook constant-overlap-add setting.

Quick start

  1. Select exactly one Sound. Mono and multichannel sources are accepted, but the output is always mono.
  2. Run Partial_Editing_Resynthesis.praat.
  3. Start with Clean Texture Resynth or Dense Partials.
  4. Use Max_partials_per_frame to move between sparse and rich reconstructions.
  5. Use Transpose_semitones for musical transposition and Additional_frequency_scale for an additional multiplicative spectral scale.
  6. Choose stable jitter for continuity or frame-random jitter for the older grainier texture.

Presets

PresetKey overrides
CustomNo override.
Clean Texture ResynthFreq jitter 0.5 Hz; amp jitter 0.02; 20 partials.
Diffuse Texture8 Hz; 0.3; 15 partials.
Sparse Partials5 partials; 2 Hz; 0.1.
Dense Partials30 partials; 1 Hz; 0.05.
Pitch Up Octave+12 st; 1 Hz; 0.05.
Pitch Down Octave−12 st; 1 Hz; 0.05.
Spectral Scale UpAdditional frequency scale 1.5; 2 Hz; 0.1.
Spectral Scale DownAdditional frequency scale 0.7; 2 Hz; 0.1.
Glassy Shimmer15 Hz; 0.4; 20 partials; max frequency 12 kHz (clipped to Nyquist if necessary).
RoboticNo frequency or amplitude jitter; 12 partials.
Whisper Ghost4 partials; 10 Hz; 0.5; max frequency 6 kHz.

Controls

ControlDefaultMeaning
Window_length0.060 sDuration of each Hann analysis/resynthesis frame.
Hop_size0.015 sTime between frame starts. The final frame is shifted when necessary so the source end is covered.
Min_frequency / Max_frequency60 / 8000 HzPeak-search band. Max is clipped to source Nyquist.
Max_partials_per_frame15Maximum number of strongest peaks resynthesized in each frame.
Freq_jitter3 HzMaximum detuning amount around each detected FFT-bin frequency; negative values are clamped to 0.
Amp_jitter0.1Relative amplitude variation; internally clamped to 0…1.
Jitter_behaviourStable per-frequencyStable mode deterministically maps each FFT bin to a fixed detune/gain offset. Frame-random mode generates new random offsets in each frame.
Transpose_semitones0Frequency multiplier 2^(st/12).
Additional_frequency_scale1.0Additional positive multiplier applied to every resynthesized partial frequency.

The final partial frequency is approximately (detected_frequency + jitter) × 2^(transpose/12) × additional_scale. Components below 20 Hz or at/above working Nyquist are omitted from the synthesized grain.

Processing pipeline

  1. For multichannel input, measure each source channel and choose the highest-RMS channel as the analysis driver. This avoids cancellation that could occur if anti-phase channels were averaged.
  2. Create a mono output buffer and a separate overlap-add window-sum buffer at the original sample rate and duration.
  3. For every frame: extract a Hann-windowed segment, compute its Spectrum and Ltas, find strong peaks one by one, and suppress ±40 Hz around each chosen peak before searching again.
  4. Apply the selected frequency/amplitude jitter, transposition and additional frequency scaling.
  5. Synthesize all accepted partials for that frame as sine waves under an analytic Hann window.
  6. Overlap-add every grain into the output.
  7. Divide by the exact accumulated Hann-window envelope.
  8. If the result is non-silent, apply Scale intensity: 70.

Channels, duration, randomness and level

Output name: <source>_resynth_<preset>.

Visualization

The figure shows original and resynthesized waveforms on one shared amplitude scale, original/output spectrograms, and a summary of partial count, jitter, transposition and frequency scale.

Multichannel display detail: analysis uses the highest-RMS channel, but the “Original Spectrogram” visualization currently displays source channel 1. On multichannel material these can therefore represent different source channels.

Historical, technological and compositional context

Sinusoidal analysis/resynthesis became a major digital sound-modeling technique in the 1980s. McAulay and Quatieri's 1986 sinusoidal representation estimated short-time sinusoidal frequencies, amplitudes and phases from spectral peaks and tracked those components through time. In computer music, Serra and Smith's 1990 Spectral Modeling Synthesis extended spectral modeling by separating a sound into deterministic sinusoidal trajectories and a stochastic residual.

This Praat tool deliberately takes a more reduced, compositional approach. It keeps only a limited set of strong peaks in each frame, does not track partial trajectories across frames, and does not resynthesize a noise residual. Those omissions are not merely technical limitations: they are what make sparse, robotic, diffuse and ghost-like textures possible. The source is treated as material for a new additive reconstruction rather than as something that must be reproduced transparently.

Further reading

McAulay, R. J., & Quatieri, T. F. (1986). “Speech Analysis/Synthesis Based on a Sinusoidal Representation.” IEEE Transactions on Acoustics, Speech, and Signal Processing, 34(4), 744–754.

Serra, X., & Smith, J. (1990). “Spectral Modeling Synthesis: A Sound Analysis/Synthesis Based on a Deterministic plus Stochastic Decomposition.” Computer Music Journal, 14(4), 12–24.