Creative Formant Manipulations — User Guide

Static spectral-envelope transformation guided by robust formant landmarks. The tool measures median formant locations, maps selected landmarks to new targets, and reshapes the original spectrum while preserving its complex phase.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 2.3 (2026) License: MIT License Repo: Praat AudioTools
Contents:

What this does

Creative Formant Manipulations v2.3 is a spectral-envelope landmark processor. Praat's FormantPath analysis is used only to estimate robust median formant locations. Those landmarks guide a smooth, time-invariant frequency-domain reshaping of the original Sound.

The processor does not resynthesize the voice from LPC data and does not filter the Sound through a FormantGrid. Instead, each channel is transformed to a complex spectrum, a smooth real-valued gain curve is applied to its magnitude, and the spectrum is transformed back to audio. Because the same real gain is applied to the real and imaginary components, the original spectral phase is preserved bin by bin.

Three current manipulations: Global formant shift moves all reliable landmarks by one ratio; F2 focus moves only F2; Formant spacing expands or contracts the spacing of the landmarks around F2 (or F1 if F2 is unavailable).

Quick start

  1. Select exactly one Sound object.
  2. Choose a preset, or choose Manual.
  3. For Manual, choose Global formant shift, F2 focus, or Formant spacing and set its ratio.
  4. Use Strength dB to control how strongly energy is redistributed between the measured and target regions.
  5. Set Dry/wet mix and the desired output-level mode.
  6. Run the script. The processed object is named <source>_CFM2_<preset>.
Input requirement: the Sound must be at least 200 ms long and must provide at least two reliable formant landmarks. The analysis is most meaningful for material with stable resonant spectral structure.

Presets

PresetManipulationValues overridden
ManualUses the form settingsNo preset overrides
Vocal LiftGlobal formant shiftGlobal ratio 1.22; strength 12 dB; dry/wet 1.0
Giant DarkGlobal formant shiftGlobal ratio 0.72; strength 18 dB; dry/wet 1.0
F2 LaserF2 focusF2 ratio 1.75; strength 24 dB; dry/wet 1.0
Wide AlienFormant spacingSpacing factor 1.65; strength 20 dB; dry/wet 1.0
Compact VowelFormant spacingSpacing factor 0.62; strength 18 dB; dry/wet 1.0

Presets do not change Max formant Hz, output-level mode, ceiling peak, visualization, or playback settings.

Controls

ControlDefaultWhat it controls
Manipulation typeGlobal formant shiftSelects the landmark mapping used in Manual mode.
Max formant Hz5500Requested FormantPath ceiling. It is automatically reduced when necessary to stay safely below Nyquist.
Global ratio1.30Multiplies every reliable formant landmark. Values above 1 move targets upward; values below 1 move them downward.
F2 ratio1.45Moves F2 only; the other measured landmarks remain at their original positions.
Spacing factor1.30Expands or contracts landmark distances around F2. F2 itself remains the pivot. If F2 is unavailable, F1 is used as the pivot.
Strength dB15Controls the maximum spectral redistribution. Valid range is greater than 0 and at most 36 dB.
Dry/wet mix1.0Linear amplitude blend: 0 = dry bypass, 1 = fully processed.
Output level modeNatural levelNatural level, Safety ceiling, or Peak normalize.
Ceiling peak0.95Target used by Safety ceiling and Peak normalize; must be greater than 0 and at most 1.
Draw visualizationYesDraws waveform, landmark, spectrogram, and summary panels.
Play resultYesPlays the result after processing.

Fixed analysis settings

v2.3 keeps the core analysis settings internal rather than exposing them in the form: 5 ms time step, 30 ms window, up to five formants, and pre-emphasis from 35 Hz. The final static landmark for each formant is the median of its FormantPath-derived track.

How the sound is processed

  1. Analysis signal: multichannel input is folded to mono for landmark estimation. If that fold nearly cancels, the script automatically analyzes the real input channel with the highest RMS instead.
  2. Landmark extraction: FormantPath/Burg analysis estimates up to five formants. A native median query supplies one robust static frequency landmark for each formant.
  3. Target mapping: the chosen manipulation maps measured landmarks to target frequencies. Targets are constrained to 80 Hz through Nyquist minus 80 Hz.
  4. Spectral redistribution: for every moved landmark, the script creates a broad Gaussian dip around the measured frequency and a matching broad lift around the target. The sum of all moves is limited to ±Strength dB.
  5. Whole-file FFT: the same static gain curve is applied independently to each channel's complex spectrum. This is a time-invariant transformation: there is no frame-by-frame formant trajectory, LFO, freezing, scrambling, or temporal crossfade.
  6. Dry/wet: the processed and original channel are mixed sample-for-sample using a linear amplitude crossfade.
Region width: each moved landmark uses a broad frequency region whose width depends on formant index and measured frequency, with an upper cap of 900 Hz. These widths describe the spectral-envelope shaping regions; they are not LPC resonator bandwidths.
No randomness: v2.3 contains no random permutation or random seed. Re-running the same settings on the same input produces the same landmark mapping and spectral gain curve.

Channels & timing

The landmark analysis is shared, but the spectral processing is performed independently on every input channel. Every channel receives the same frequency-domain gain curve, so mono, stereo, and higher channel counts are preserved.

PropertyBehavior
Channel countPreserved.
Sample ratePreserved.
DurationRestored to the original duration after inverse FFT padding.
Start timeThe work copy is shifted to 0 for processing; the final output is shifted back to the source xmin.
PhaseThe original complex spectral phase is preserved bin by bin; only magnitude is multiplied by the real gain curve.

Output level & playback

ModeBehavior
Natural levelNo final level scaling is applied.
Safety ceilingScales down only when the output peak exceeds Ceiling peak. Quieter output is left unchanged.
Peak normalizeIf the output is non-silent, scales its peak to Ceiling peak, upward or downward as required.

If Natural level leaves the stored result above 1.0, playback uses a temporary copy scaled to 0.95. The stored output itself is not altered by this playback safeguard.

A true bypass occurs before analysis when Dry/wet is 0 or the selected manipulation ratio is exactly neutral. The requested output-level mode is still applied to that bypass copy.

Visualization

The Picture window is diagnostic rather than a second processing stage. It contains:

For Sounds longer than 8 seconds, waveform and spectrogram drawing use a central 8-second excerpt so visualization cost does not grow with the full recording. The actual audio processing still covers the complete Sound.

Important: the landmark panel shows the static median landmarks that control the spectral-envelope mapping. It does not represent time-varying formant tracks.

Further Reading