Jitter-Shimmer Formant Mapping — User Guide

Measures whole-sound jitter and shimmer, then uses those measurements to control a static spectral-envelope remapping around five formant landmarks.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 3.1.1 (2026) License: MIT License Repo: Praat AudioTools
Contents:

What this does

Jitter-Shimmer Formant Mapping converts two voice-quality measurements into two static spectral-envelope transformations. Shimmer controls the target positions of F1–F2; jitter controls F3–F5. The five formants are analysis landmarks only: the script does not rebuild an LPC filter and does not move formant tracks frame by frame.

The transformation is applied to the complex spectrum of each output channel as a smooth real-valued gain curve. Because the same gain is applied to the real and imaginary components at each frequency, the original spectral phase is preserved.

Important: jitter and shimmer are measured once over the whole analysis sound. The resulting formant ratios are therefore static for the entire file. This is not a time-varying jitter/shimmer follower.

Signal flow

  1. Choose an analysis source. Mono input uses itself. For multichannel input, the channel with the highest RMS is used for analysis; the output channel structure is not collapsed.
  2. Estimate pitch range. With auto detection enabled, a wide pitch analysis supplies 10th/90th-percentile estimates used to set a practical floor and ceiling.
  3. Measure voice perturbation. Praat's voice-analysis path provides whole-sound Jitter (local) and Shimmer (local). If a reliable measurement cannot be obtained, that measurement falls back to zero rather than being invented.
  4. Find five spectral landmarks. A FormantPath analysis supplies median F1–F5 frequencies. These are static reference frequencies, not synthesis poles.
  5. Calculate target landmarks. Shimmer determines the F1–F2 ratio; jitter determines the F3–F5 ratio.
  6. Warp the spectral envelope. Broad Gaussian gain regions reduce energy around each measured landmark and add energy around its target position. The accumulated gain is limited by Envelope_strength_dB.
  7. Optional pitch stage. If enabled and the selected preset uses a pitch multiplier other than 1.0, Praat Manipulation resynthesis changes pitch after the spectral warp.
  8. Dry/wet and output level. Each channel is mixed with its own dry channel, then the selected output-level policy is applied.

Quick start

  1. Select one or more Sound objects in Praat.
  2. Run Jitter-Shimmer_Formant_Mapping.praat.
  3. Start with Modal (Subtle) to hear a restrained mapping.
  4. Use Global intensity to scale the measured jitter/shimmer before they drive the mapping.
  5. Use Dry/wet mix to blend the mapped result with the original.
  6. Enable Apply pitch shift only if you also want the preset's separate pitch multiplier.
  7. Enable Advanced settings for pitch range, confidence, envelope width/strength and output-level policy.
Exact bypass: Dry_wet_mix = 0 returns a sample-identical copy of the original Sound, preserving channels, duration, sample rate and start time. Output-level processing is deliberately skipped in this bypass path.

Presets

PresetF1–F2 weightF3–F5 weightEnvelopePitch multiplier*
Modal (Subtle)+0.10+0.1010 dB, width ×1.101.00
Breathy (Brighter)+0.50+0.3016 dB, width ×1.251.08
Creaky (Darker)−0.40−0.2015 dB, width ×1.150.92
Tense (Sharp)+0.20+0.6019 dB, width ×0.851.05
Relaxed (Smooth)+0.30−0.2012 dB, width ×1.450.98
CustomShimmer_to_F1F2Jitter_to_F3F5Advanced values1.00

*The pitch multiplier is used only when Apply pitch shift is enabled. Modal and Custom use 1.00, so enabling the pitch stage does not change pitch in those two presets.

Preset override rule: presets 1–5 replace the two mapping weights and also replace Envelope_strength_dB and Envelope_width_scale. Those four user-entered values are fully user-controlled only in Custom. Other advanced settings remain active for all presets.

Controls

Main form

ControlDefaultMeaning
PresetModalSelects one of five fixed mappings or Custom.
Global_intensity1.0Scales the measured jitter and shimmer before logarithmic normalization. Larger values create larger target shifts, subject to the final ratio clamp.
Shimmer_to_F1F20.3Custom-mode mapping weight for F1–F2. Positive values move targets upward; negative values move them downward.
Jitter_to_F3F50.3Custom-mode mapping weight for F3–F5. Positive values move targets upward; negative values move them downward.
Auto_detect_pitch_rangeOnAdapts the voice-analysis pitch floor/ceiling from a preliminary pitch estimate.
Apply_pitch_shiftOffEnables the preset-dependent, separate pitch stage after the spectral warp.
Dry_wet_mix1.00 = exact dry bypass; 1 = fully processed. Values outside 0–1 are clamped.
Advanced_settingsOffOpens the second settings form.
Draw_visualizationOnDraws analysis for the first selected Sound only.
Play_resultOnPlays each produced result after processing.

Advanced settings

ControlDefaultMeaning
Manual pitch floor / ceiling75 / 600 HzUsed directly when auto range is off and as fallback limits if auto estimation is not usable.
Keep intermediatesOffKeeps the final Pitch, pulse PointProcess and extracted Formant landmark object for each processed Sound.
Max_formant_hz5500 HzUpper reference for FormantPath analysis; constrained internally against Nyquist.
Envelope_width_scale1.0Scales the broad Gaussian regions used for redistribution. Internally clamped to 0.2–4.0. Presets 1–5 override it.
Envelope_strength_dB15 dBMaximum magnitude of the accumulated spectral gain curve. Internally clamped to 0.1–36 dB. Presets 1–5 override it.
Require_formant_confidenceOnBypasses the spectral warp if fewer than two usable landmarks are found or their total span is under 600 Hz.
Output_level_modeNatural levelNatural / Safety ceiling / Peak normalize.
Ceiling_peak0.95Target used by Safety ceiling and Peak normalize. Invalid values fall back to 0.95.

How the mapping works

1. Whole-sound jitter and shimmer

The script uses Praat's voice-analysis chain: Sound → Pitch (cc) → PointProcess (cc) → Voice report, then reads Jitter (local) and Shimmer (local). The time range passed to Voice report is 0–0, so the values summarize the whole analysis Sound.

To make the control response less explosive at high perturbation values, the measured percentages are logarithmically compressed. With Global intensity = 1, roughly 5% jitter maps to a normalized drive of 1, and roughly 10% shimmer maps to 1. The resulting F1–F2 and F3–F5 frequency ratios are finally constrained to 0.5–2.0.

2. Static formant landmarks

The script runs FormantPath (Burg) on the analysis channel and takes the median frequency of F1 through F5 over the whole Sound. These medians are the five reference landmarks used to construct the static spectral map.

3. Spectral redistribution

For every active landmark, the gain curve contains a broad negative Gaussian centred at the measured frequency and a matching positive Gaussian centred at its target. The five contributions are summed, limited to ±Envelope_strength_dB, converted from dB to linear gain, and multiplied into the complex Spectrum.

This is an energy redistribution around landmark regions, not literal movement of LPC resonances. A target can only emphasize energy already represented in the broadband spectrum; it does not synthesize a new oscillator at that frequency.

4. Optional pitch shift

The pitch stage is independent of the jitter/shimmer mapping. When enabled, the script creates a Manipulation object, multiplies its PitchTier by the preset multiplier, and resynthesizes with overlap-add. The spectral warp has already been completed before this stage.

Input, output & level behavior

Processed naming: original_JitterShimmer_PresetName. Exact bypass uses original_JitterShimmer_Bypass; silent/near-silent analysis falls back to original_JitterShimmer_SilenceBypass.

Visualization

Visualization is drawn for the first selected Sound only and does not alter the audio result. Tests with visualization on and off produced sample-identical output.

For Sounds longer than five seconds, the spectrogram comparison shows the first five seconds, while the jitter/shimmer and landmark calculations still refer to the full analysis Sound.

Notes & limitations

Further reading