Creative Formant Manipulations — User Guide
Static spectral-envelope transformation guided by robust formant landmarks. The tool measures median formant locations, maps selected landmarks to new targets, and reshapes the original spectrum while preserving its complex phase.
What this does
Creative Formant Manipulations v2.3 is a spectral-envelope landmark processor. Praat's FormantPath analysis is used only to estimate robust median formant locations. Those landmarks guide a smooth, time-invariant frequency-domain reshaping of the original Sound.
The processor does not resynthesize the voice from LPC data and does not filter the Sound through a FormantGrid. Instead, each channel is transformed to a complex spectrum, a smooth real-valued gain curve is applied to its magnitude, and the spectrum is transformed back to audio. Because the same real gain is applied to the real and imaginary components, the original spectral phase is preserved bin by bin.
Quick start
- Select exactly one Sound object.
- Choose a preset, or choose Manual.
- For Manual, choose Global formant shift, F2 focus, or Formant spacing and set its ratio.
- Use Strength dB to control how strongly energy is redistributed between the measured and target regions.
- Set Dry/wet mix and the desired output-level mode.
- Run the script. The processed object is named
<source>_CFM2_<preset>.
Presets
| Preset | Manipulation | Values overridden |
|---|---|---|
| Manual | Uses the form settings | No preset overrides |
| Vocal Lift | Global formant shift | Global ratio 1.22; strength 12 dB; dry/wet 1.0 |
| Giant Dark | Global formant shift | Global ratio 0.72; strength 18 dB; dry/wet 1.0 |
| F2 Laser | F2 focus | F2 ratio 1.75; strength 24 dB; dry/wet 1.0 |
| Wide Alien | Formant spacing | Spacing factor 1.65; strength 20 dB; dry/wet 1.0 |
| Compact Vowel | Formant spacing | Spacing factor 0.62; strength 18 dB; dry/wet 1.0 |
Presets do not change Max formant Hz, output-level mode, ceiling peak, visualization, or playback settings.
Controls
| Control | Default | What it controls |
|---|---|---|
| Manipulation type | Global formant shift | Selects the landmark mapping used in Manual mode. |
| Max formant Hz | 5500 | Requested FormantPath ceiling. It is automatically reduced when necessary to stay safely below Nyquist. |
| Global ratio | 1.30 | Multiplies every reliable formant landmark. Values above 1 move targets upward; values below 1 move them downward. |
| F2 ratio | 1.45 | Moves F2 only; the other measured landmarks remain at their original positions. |
| Spacing factor | 1.30 | Expands or contracts landmark distances around F2. F2 itself remains the pivot. If F2 is unavailable, F1 is used as the pivot. |
| Strength dB | 15 | Controls the maximum spectral redistribution. Valid range is greater than 0 and at most 36 dB. |
| Dry/wet mix | 1.0 | Linear amplitude blend: 0 = dry bypass, 1 = fully processed. |
| Output level mode | Natural level | Natural level, Safety ceiling, or Peak normalize. |
| Ceiling peak | 0.95 | Target used by Safety ceiling and Peak normalize; must be greater than 0 and at most 1. |
| Draw visualization | Yes | Draws waveform, landmark, spectrogram, and summary panels. |
| Play result | Yes | Plays the result after processing. |
Fixed analysis settings
v2.3 keeps the core analysis settings internal rather than exposing them in the form: 5 ms time step, 30 ms window, up to five formants, and pre-emphasis from 35 Hz. The final static landmark for each formant is the median of its FormantPath-derived track.
How the sound is processed
- Analysis signal: multichannel input is folded to mono for landmark estimation. If that fold nearly cancels, the script automatically analyzes the real input channel with the highest RMS instead.
- Landmark extraction: FormantPath/Burg analysis estimates up to five formants. A native median query supplies one robust static frequency landmark for each formant.
- Target mapping: the chosen manipulation maps measured landmarks to target frequencies. Targets are constrained to 80 Hz through Nyquist minus 80 Hz.
- Spectral redistribution: for every moved landmark, the script creates a broad Gaussian dip around the measured frequency and a matching broad lift around the target. The sum of all moves is limited to ±Strength dB.
- Whole-file FFT: the same static gain curve is applied independently to each channel's complex spectrum. This is a time-invariant transformation: there is no frame-by-frame formant trajectory, LFO, freezing, scrambling, or temporal crossfade.
- Dry/wet: the processed and original channel are mixed sample-for-sample using a linear amplitude crossfade.
Channels & timing
The landmark analysis is shared, but the spectral processing is performed independently on every input channel. Every channel receives the same frequency-domain gain curve, so mono, stereo, and higher channel counts are preserved.
| Property | Behavior |
|---|---|
| Channel count | Preserved. |
| Sample rate | Preserved. |
| Duration | Restored to the original duration after inverse FFT padding. |
| Start time | The work copy is shifted to 0 for processing; the final output is shifted back to the source xmin. |
| Phase | The original complex spectral phase is preserved bin by bin; only magnitude is multiplied by the real gain curve. |
Output level & playback
| Mode | Behavior |
|---|---|
| Natural level | No final level scaling is applied. |
| Safety ceiling | Scales down only when the output peak exceeds Ceiling peak. Quieter output is left unchanged. |
| Peak normalize | If the output is non-silent, scales its peak to Ceiling peak, upward or downward as required. |
If Natural level leaves the stored result above 1.0, playback uses a temporary copy scaled to 0.95. The stored output itself is not altered by this playback safeguard.
A true bypass occurs before analysis when Dry/wet is 0 or the selected manipulation ratio is exactly neutral. The requested output-level mode is still applied to that bypass copy.
Visualization
The Picture window is diagnostic rather than a second processing stage. It contains:
- A — Original waveform and B — Processed waveform, using channel 1 and one shared amplitude scale.
- C — Spectral-envelope landmark map, with measured frequencies in grey, target frequencies in colour, and connectors showing each move.
- D/E — Original and processed spectrograms for channel 1.
- Summary strip with the static FFT method, landmark count, input/output peak, duration, channel count, and sample rate.
For Sounds longer than 8 seconds, waveform and spectrogram drawing use a central 8-second excerpt so visualization cost does not grow with the full recording. The actual audio processing still covers the complete Sound.
Further Reading
- Fant, G. (1960). Acoustic Theory of Speech Production: With Calculations Based on X-Ray Studies of Russian Articulations. Mouton. — foundational source-filter and formant theory.
- Praat manual: FormantPath. — primary documentation for the analysis object used to obtain alternative formant analyses and a preferred path.
- Oppenheim, A. V., & Schafer, R. W. (2009). Discrete-Time Signal Processing (3rd ed.). Pearson. — discrete Fourier analysis and frequency-domain filtering.