Formant Synthesis — User Guide
A KlattGrid source–filter vowel/voice synthesizer. Pitch, voicing, aspiration and breathiness form the source; four oral resonances F1–F4 form the filter. Presets configure vowel targets and synthetic voice variants, with optional vibrato and stereo rendering.
What this does
Formant Synthesis uses Praat's KlattGrid as an explicit source–filter synthesizer. The source is defined by F0, voicing amplitude, aspiration and breathiness. That source then passes through four oral resonances whose center frequencies and bandwidths are set by F1–F4 and BW1–BW4.
pitch + phonation / aspiration / breathiness → KlattGrid oral resonances F1–F4 → optional stereo field → short edge fade → optional final peak normalization
The four formants are filter resonances, not four sine oscillators. This distinction is central to the current implementation.
The script is designed for controllable vowel-like and synthetic voice timbres. It is not a complete speech synthesizer: the resonances are static during each rendered sound, there are no consonant gestures, and the model uses only four oral resonances with a compact set of source controls.
Source and filter
Source
KlattGrid provides separate tiers for the periodic voice source and for noise-related source components. In this script the relevant controls are:
- Pitch: the nominal F0, optionally modulated by vibrato.
- Voicing amplitude: periodic/phonated source level in dB.
- Aspiration amplitude: aspiration-noise source level in dB.
- Breathiness amplitude: breathiness-noise source level in dB.
All active source amplitudes receive the same short attack and release structure. The attack is at most 25 ms and the release at most 60 ms, with both automatically shortened for very short sounds.
Filter
The filter consists of four static oral resonances. Each resonance has a center frequency and a bandwidth:
source → [F1/BW1] → [F2/BW2] → [F3/BW3] → [F4/BW4] → output
Within one rendered sound, F1–F4 and their bandwidths remain constant. Presets change these targets before synthesis; they do not create time-varying vowel transitions.
Quick start
- Run
Formant_Synthesis.praat. No input Sound is required. - Choose a preset such as Vowel A, Vowel I, Aspiration-Dominant Whisper, or Extreme Formant Voice.
- Set Duration, Pitch, Sample rate, and Spatial mode.
- Use Edit formant details when you want direct F1–F4 and bandwidth control.
- Use Edit source details for voicing/noise levels, vibrato, random seed, and edge fade.
- Leave Normalize output on for target peak normalization to 0.90, or turn it off to retain the raw rendered level.
Presets
Presets load source and/or resonance settings before the optional advanced pages open. They do not change Spatial mode, Sample rate, Normalize output, Draw visualization, or Play result.
| Preset | Pitch | F1 / F2 / F3 / F4 (Hz) | BW1 / BW2 / BW3 / BW4 (Hz) | Source emphasis |
|---|---|---|---|---|
| Vowel A | form value | 730 / 1090 / 2440 / 3500 | 40 / 60 / 100 / 120 | Voice 60 dB; aspiration 8; breathiness 4 |
| Vowel E | form value | 530 / 1840 / 2480 / 3500 | 45 / 65 / 105 / 125 | Voice 60; aspiration 8; breathiness 4 |
| Vowel I | form value | 270 / 2290 / 3010 / 3500 | 35 / 70 / 110 / 130 | Voice 60; aspiration 7; breathiness 3 |
| Vowel O | form value | 570 / 840 / 2410 / 3500 | 50 / 60 / 100 / 120 | Voice 60; aspiration 8; breathiness 4 |
| Vowel U | form value | 300 / 870 / 2240 / 3500 | 40 / 55 / 95 / 115 | Voice 60; aspiration 8; breathiness 4 |
| High-F0 A | 260 Hz | 800 / 1150 / 2900 / 3900 | 35 / 50 / 90 / 110 | Voice 62; light noise; vibrato depth 0.55 st |
| Low-F0 O | 80 Hz | 450 / 800 / 2830 / 3500 | 60 / 70 / 120 / 140 | Voice 61; light noise; vibrato depth 0.35 st |
| High-F0 E | 300 Hz | 600 / 2000 / 2600 / 3800 | 30 / 55 / 95 / 115 | Voice 60; light noise; vibrato depth 0.45 st |
| Narrow-Band Synthetic Voice | form value | 400 / 1200 / 2400 / 3200 | 20 / 30 / 40 / 50 | Voice 64; no aspiration/breathiness; vibrato off |
| Aspiration-Dominant Whisper | form value | 500 / 1500 / 2500 / 3500 | 80 / 100 / 150 / 200 | Voicing 0; aspiration 72; breathiness 0; vibrato off |
| Vibrato Vocal Tone | 220 Hz | 600 / 1200 / 2400 / 3600 | 35 / 55 / 95 / 115 | Voice 62; vibrato 5.5 Hz / 0.95 st |
| Extreme Formant Voice | 180 Hz | 200 / 3000 / 4000 / 5000 | 15 / 25 / 35 / 45 | Voice 64; light noise; vibrato off |
Formant controls
Edit formant details opens a second page containing F1–F4 and BW1–BW4. The script requires:
0 < F1 < F2 < F3 < F4 all bandwidths > 0
Sample-rate headroom
The practical upper synthesis limit is 0.45 × sample rate. If F4 lies above that value, the script does not clip F4 alone. Instead it computes one common scale factor and applies it to:
- F1, F2, F3, F4
- BW1, BW2, BW3, BW4
This preserves the relative formant geometry and bandwidth proportions as much as possible. The Info window and visualization report the effective formants and the applied scale.
Source controls
| Control | Default | Behavior |
|---|---|---|
| Voicing amplitude | 60 dB | Periodic source level. A value of 0 disables the voiced source tier. |
| Aspiration amplitude | 0 dB | Aspiration-noise source level. |
| Breathiness amplitude | 0 dB | Breathiness-noise source level. |
| Enable vibrato | yes | Turns pitch modulation on/off. |
| Vibrato rate | 6 Hz | Rate of the F0 modulation. |
| Vibrato depth | 0.75 semitones | Peak excursion in semitones. |
| Random seed | 0 | 0 = unpredictable stochastic source realization; positive integer = reproducible stochastic components. |
| Edge fade | 0.02 s | Final linear click-protection fade at the beginning and end, capped at 20% of duration. |
Vibrato
Vibrato is defined in semitones, then converted multiplicatively to frequency:
F0(t) = baseF0 × 2^[(depth × sin(2π × rate × t)) / 12]
The pitch tier is sampled at approximately eight control points per vibrato cycle. This keeps the same musical pitch depth meaningful at different F0 values instead of using a fixed Hz excursion.
Random seed
A positive seed makes KlattGrid's stochastic source components reproducible for the same settings. After synthesis, the script restores Praat's unpredictable random initialization. In Detuned Stereo Chorus, both channels use the same source/filter parameter structure but are synthesized as separate voices at slightly different F0 values.
Spatial modes
| Mode | Implementation | Channels |
|---|---|---|
| Mono | One complete KlattGrid voice. | 1 |
| Stereo Micro-Delay | Left = original / √2. Right = delayed copy / √2, with delay min(8 ms, 4% of duration). | 2 |
| Detuned Stereo Chorus | Two complete KlattGrid voices: left at −5 cents, right at +5 cents. Both retain the same vibrato, formants, source envelopes, aspiration and breathiness settings. | 2 |
Stereo Micro-Delay preserves the full vowel spectrum in both channels; it does not split low and high formants between the ears. The right channel is delayed inside the same total duration, so the very end of the delayed copy is not extended beyond the requested output duration.
Detuned Stereo Chorus is true dual-voice rendering rather than a pitch-shifted copy after synthesis.
Output, duration, sample rate, and level
| Property | Current behavior |
|---|---|
| Duration | Requested Duration, up to 120 seconds. |
| Sample rate | 8–192 kHz. KlattGrid output is resampled to the requested rate when necessary using Praat's sinc resampling with precision 50. |
| Channels | Mono in Mono mode; stereo in Micro-Delay and Chorus. |
| Output name | formant_<preset name> with spaces replaced by underscores. |
| Save to file | Not part of the current script. The output remains as a Sound object in Praat. |
Normalization
The final output is measured after spatialization and the outer edge fade. If Normalize output is enabled and the Sound is non-zero, the script runs:
Scale peak: 0.90
This is target peak normalization: quieter non-zero results are raised as well as louder results being reduced. With Normalize output off, no final peak scaling is applied.
Visualization and QC
| Panel | What it actually shows |
|---|---|
| A — F1–F2 Vowel Space | Reference points for /i e a o u/ plus the current effective F1/F2 target when it falls inside the displayed range. |
| B — KlattGrid Pitch Control | The exact center-voice F0 control trajectory used by the synthesis model. In Chorus mode, ±5-cent reference lines are also shown. |
| C — Model → Measurement | Measured spectrogram of a representative final-output channel with the four effective synthesis-resonance targets overlaid. |
| D — Measured Output Spectrum | Measured spectrum of the same representative channel with F1–F4 target markers. |
For stereo output, the representative channel is whichever channel has the higher whole-file RMS.
The QC strip summarizes source levels and seed, effective F1–F4 values and Nyquist scale, nominal pitch/vibrato range, spatial mode, and pre/post-normalization peak/RMS measurements.
Further Reading
- Praat Manual — KlattGrid. Documents KlattGrid as a time-varying source–filter model with separate source, vocal-tract filter, and source/filter-interaction tiers.
- Praat Manual — Source-filter synthesis. Introduces the source–filter model and Praat's synthesis workflow.
- Klatt, D. H. & Klatt, L. C. (1990). “Analysis, synthesis, and perception of voice quality variations among female and male talkers.” Journal of the Acoustical Society of America, 87, 820–857. DOI: 10.1121/1.398894.