Individual Formant Stretcher — User Guide
Static spectral-envelope transformation that moves individually measured formant-landmark regions by independent semitone amounts while preserving the original complex-spectrum phase.
What this does
Individual Formant Stretcher analyses a selected Sound and estimates up to five robust, time-aggregated formant landmarks. Each landmark can then be moved independently by a number of semitones, with an optional global transpose added to all five.
The landmarks are used as spectral-envelope control points. The script does not move Praat Formant objects through time, does not inverse-filter the source, and does not perform LPC or FormantGrid resynthesis. Instead, it builds one static spectral gain curve for the whole sound: energy is attenuated around the measured landmark region and enhanced around its transposed destination. That real-valued gain curve is applied equally to the real and imaginary components of the complex spectrum, preserving the original spectral phase.
Quick start
- Select exactly one Sound object in Praat.
- Run
Individual_Formant_Stretcher.praat. - Choose a preset, or use Custom and enter independent semitone shifts for F1–F5 plus an optional global shift.
- Set Bandwidth scale to control how broad each affected spectral region is, and Strength dB to control the maximum spectral shaping depth.
- For multichannel material, choose which signal is used for the landmark analysis. The resulting map is then applied independently to every channel.
- Choose Dry/Wet and the desired output-level mode, then run the script.
Processing pipeline
Spectral remapping
For every active landmark, the script creates two broad Gaussian regions: a cut centred on the measured landmark and a boost centred on the target. The per-landmark contribution is approximately:
The contributions of all active landmarks are summed, then the total spectral gain is limited to ±Strength dB. The linear multiplier is 10^(gain_dB/20).
Region width
Bandwidth scale does not scale LPC formant bandwidths. It scales the width of the Gaussian spectral regions used for the remapping. Before scaling, each width is the larger of 24% of the measured landmark frequency or a formant-specific minimum:
| Landmark | Minimum base width |
|---|---|
| F1 | 180 Hz |
| F2 | 260 Hz |
| F3 | 360 Hz |
| F4 | 450 Hz |
| F5 | 550 Hz |
After applying Bandwidth scale, the working width is constrained to 90–1400 Hz.
Presets
Preset transpose values are additive with the preset's Global value. For example, a preset with Global +2 and F2 +2 sends F2 upward by a total of +4 semitones. Presets do not override Analysis source, confidence, Dry/Wet, output-level mode, ceiling, visualization, or playback.
| Preset | F1 / F2 / F3 / F4 / F5 | Global | Width | Strength |
|---|---|---|---|---|
| Natural (no change) | 0 / 0 / 0 / 0 / 0 st | 0 st | User value; irrelevant on bypass | 18 dB |
| Compress Vowel Space | +3 / +2 / 0 / −2 / −3 st | 0 st | User value | 14 dB |
| Expand Vowel Space | −3 / −2 / 0 / +2 / +3 st | 0 st | User value | 14 dB |
| Brighten Spectrum | 0 / +4 / +6 / +7 / +8 st | 0 st | 0.70× | 18 dB |
| Darken Spectrum | 0 / −4 / −6 / −7 / −8 st | 0 st | 1.40× | 18 dB |
| Male to Female | −1 / +2 / +3 / +3 / +2 st | +2 st | 0.85× | 16 dB |
| Female to Male | +1 / −2 / −3 / −3 / −2 st | −2 st | 1.15× | 16 dB |
| Robot Voice (harmonic) | 0 / +12 / +19 / +24 / +28 st | 0 st | 0.50× | 24 dB |
| Alien Creature | +8 / −5 / +12 / −8 / +15 st | 0 st | 1.50× | 22 dB |
| Demon Voice | −7 / −12 / −8 / −15 / −10 st | −5 st | 2.00× | 24 dB |
| Chipmunk Extreme | +5 / +8 / +10 / +12 / +12 st | +7 st | 0.60× | 22 dB |
| Giant Extreme | −8 / −12 / −10 / −14 / −12 st | −8 st | 2.50× | 24 dB |
| Spectral Inversion | +12 / +5 / 0 / −5 / −12 st | 0 st | 1.00× | 20 dB |
| Harmonic Series | 0 / +12 / +19 / +24 / +28 st | 0 st | 0.40× | 24 dB |
| Chaos Mode | +9 / −11 / +14 / −6 / +17 st | 0 st | 1.80× | 24 dB |
Parameters
Individual landmark control
| Control | Default | Meaning |
|---|---|---|
| F1–F5 transpose semitones | 0 | Independent frequency displacement for each measured landmark. |
| Global transpose semitones | 0 | Additional displacement applied to every landmark before its individual value. |
Envelope shape
| Control | Default | Meaning |
|---|---|---|
| Bandwidth scale | 1.0 | Scales the width of the spectral regions surrounding the measured and target landmarks. Valid range in the script: greater than 0 and at most 4. |
| Strength dB | 18 dB | Sets the nominal per-landmark shaping strength and the final ±dB limit of the combined gain curve. Valid range: greater than 0 and at most 36 dB. |
Analysis
| Control | Default | Meaning |
|---|---|---|
| Max formant Hz | 5500 Hz | Middle formant ceiling supplied to FormantPath analysis. It is automatically reduced when required by the input Nyquist frequency. |
| Analysis source | Loudest channel | Channel 1, loudest channel by whole-file RMS, or mono sum. The chosen source only determines the landmarks; all original channels are processed and preserved. |
| Require formant confidence | On | Bypasses processing unless at least two valid landmarks span at least 600 Hz, reducing the chance of treating a narrow spectral-line cluster as a formant envelope. |
Output
| Control | Default | Meaning |
|---|---|---|
| Dry/Wet mix | 1.0 | 0 = unchanged dry copy; 1 = fully processed. Intermediate values mix each processed channel with its corresponding dry channel. |
| Output level mode | Natural level | Natural level / Safety ceiling / Peak normalize. |
| Ceiling peak | 0.95 | Peak target used by Safety ceiling and Peak normalize. |
| Draw visualization | On | Draw the Praat AudioTools diagnostic page. |
| Play after processing | On | Audition the result. If its natural-level peak exceeds 1.0, playback uses a temporary 0.95-peak copy while leaving the stored output unchanged. |
Landmark analysis
The script uses Praat FormantPath (Burg) only to obtain robust spectral landmarks. The fixed analysis settings are:
| Setting | Value |
|---|---|
| Time step | 5 ms |
| Maximum number of formants | 5 |
| Window length | 30 ms |
| Pre-emphasis from | 35 Hz |
| FormantPath ceiling step | 0.05 |
| Alternative analyses on each side | 4 |
After the preferred Formant is extracted from the FormantPath, each F1–F5 landmark is the median frequency over the whole file. The effect is therefore intentionally static: it does not follow instantaneous vowel changes from frame to frame.
Input & output behavior
| Property | Behavior |
|---|---|
| Input | Exactly one Sound object. |
| Channels | Preserved. One common envelope map is derived from the selected analysis source, then applied independently to every channel. |
| Duration | Preserved. |
| Sampling frequency | Preserved. |
| Start time (xmin) | Preserved. Processing is performed on a temporary zero-based copy and the original start time is restored at the end. |
| Randomness | None. |
| Processed output name | <source>_FormantStretch_<PresetName> |
| Bypass output name | <source>_FormantStretch_Bypass |
Output-level modes
| Mode | Behavior |
|---|---|
| Natural level | No hidden scaling. Spectral boosts can therefore produce peaks above 1.0. |
| Safety ceiling | Attenuates only when the processed peak exceeds Ceiling peak. It never boosts a quiet result. |
| Peak normalize | Scales every non-silent processed result to Ceiling peak. |
On an explicit bypass path, the audio is intentionally returned unchanged and output-level processing is skipped, even if Safety ceiling or Peak normalize was selected.
Visualization
When enabled, the script draws a suite-standard 8-inch diagnostic page. For files longer than 8 seconds, the waveform and spectrogram displays use a centred 8-second excerpt.
- Original / Output waveforms: channel 1, displayed with one shared amplitude scale.
- Landmark map: measured landmarks in grey, target frequencies in blue, with arrows showing active moves.
- Original / Output spectrograms: side-by-side comparison over the displayed excerpt.
- Summary strip: preset, all six transpose values, region-width scale, strength, mix, input/output peak, and processing or bypass status.
The picture is diagnostic only; visualization choices do not alter the stored Sound.
Notes & limitations
- Static map: one set of median landmarks and one gain curve are used for the entire file. Material with changing vowels or changing spectral envelopes will not receive time-varying formant tracking.
- Analysis-dependent: the effect is meaningful only when the chosen analysis source yields plausible formant-like spectral landmarks. The confidence option is designed to reject narrow line spectra rather than inventing a vocal-tract interpretation.
- Not a pitch shifter: the processor changes spectral weighting, not the frequencies of all partials by a common ratio.
- Not LPC resynthesis: Bandwidth scale is the width of broad Gaussian spectral regions, not the bandwidth of LPC poles.
- Destination clamping: extreme upward or downward settings can be limited by the 80 Hz and Nyquist−80 Hz boundaries, so the realized shift can be smaller than the requested semitone value.
- Natural-level peaks: overlapping boosts from several moved landmarks can raise the waveform peak substantially; use Safety ceiling when you need a bounded output without normalizing quieter results upward.
Further Reading
- Praat manual — FormantPath / Sound: To FormantPath (burg)... Documentation for the alternative-ceiling Burg analyses used to estimate the spectral landmarks in this tool.
- Praat manual — Sound: To Spectrum... Documentation for Praat's complex Fourier spectrum, including separate real and imaginary components and inverse reconstruction with Spectrum: To Sound.
- Fant, G. (1960). Acoustic Theory of Speech Production: With Calculations Based on X-ray Studies of Russian Articulations. Mouton. Background on formants, vocal-tract resonances, and the speech spectral envelope.