Pitch Processor — User Guide
Two complementary pitch-based processors: a stereo detune engine with PSOLA or varispeed transposition, and a time-delayed canon built from independently transposed voices.
What this does
Pitch Processor contains two independent processing modes:
- Stereo Pitch Detune creates a new stereo signal from a mono working copy. The two channels can be transposed by PSOLA (duration-preserving) or by varispeed (tape-style pitch and duration change), with optional symmetric detuning and Haas delay.
- Time-Delayed Canon creates multiple varispeed-transposed voices, gives each voice its own entry delay and intensity target, and mixes them as mono, constant-power stereo spread, or alternating left/right.
Both modes end with an optional raised-cosine fade and an attenuate-only peak ceiling. Quiet results are not normalized upward.
Quick start
- Select exactly one Sound object.
- Run
Pitch_Processor.praat. - Choose Mode: Stereo Pitch Detune or Time-Delayed Canon.
- Choose the preset for the selected mode.
- Choose Settings:
- Run the preset as it is — render immediately with the preset/default values.
- Edit main settings — open only the controls relevant to the selected mode.
- Edit main and advanced settings — open the mode controls, then output/analysis controls.
- Choose whether to draw the visualization and play the result.
- Click OK.
Interface & settings dialogs
The main form is intentionally compact. It contains only the mode, the mode-specific preset menus, the settings switch, and the draw/play toggles. Processing parameters live in optional dialogs so that Detune and Canon do not expose irrelevant controls.
Main form
| Field | Default | Meaning |
|---|---|---|
| Mode | Stereo Pitch Detune | Selects the processing engine. |
| Detune_preset | Classic detune | Used when Mode = Stereo Pitch Detune. |
| Canon_preset | Custom | Used when Mode = Time-Delayed Canon. |
| Settings | Run the preset as it is | Controls whether the optional parameter dialogs appear. |
| Draw_visualization | Yes | Draws the 8-inch analysis page in the Picture window. |
| Play_after_processing | Yes | Plays the completed result. |
0 4 7 11 and 0 3 7 10). When one of these presets is selected, that stored pattern takes precedence over the step rule unless you enter your own Semitone_list. A user Semitone_list always has highest priority and also determines the number of voices.
Mode 1 — Stereo Pitch Detune
The Detune engine creates two transposed copies of the mono working source and combines them as left and right channels. Stereo_detune_semitones specifies the interval between the two channels.
Balance
| Detune_balance | Left shift | Right shift |
|---|---|---|
| Right channel only | 0 ST | +Detune ST |
| Symmetric split | −Detune/2 | +Detune/2 |
PSOLA (duration preserving)
For a non-zero channel shift, Praat creates a Manipulation object, extracts its PitchTier, multiplies all defined F0 values by 2^(semitones/12), replaces the tier, and resynthesizes by overlap-add. The result is then resampled to Output_sample_rate. The pitch change preserves the source duration before any Haas delay is added.
Varispeed (tape transposition)
Varispeed overrides the working copy's sampling frequency by the pitch ratio, then resamples it to the requested output rate. Pitch and duration therefore change together:
active duration ≈ source duration / ratio
The two channels are padded to the same final duration before stereo combination. With unequal varispeed shifts, this padding does not restore time alignment; it only makes the channel lengths equal.
Haas delay
Haas_delay_ms prepends silence to one channel after pitch processing. Positive values delay the right channel; negative values delay the left. The absolute value is limited to 200 ms. A Haas delay therefore increases the final output duration.
The Info window also reports the nominal beat rate at the measured source mean F0 when a usable mean F0 exists.
Mode 2 — Time-Delayed Canon
The Canon engine creates one mono varispeed copy per voice. Each voice has a semitone value, optional cent jitter, entry delay, intensity target, and pan coefficients.
Pitch-pattern priority
- If Semitone_list is non-empty, it becomes the complete pitch pattern and its length becomes the voice count.
- Otherwise, if the selected preset contains an explicit pattern, that preset pattern is used.
- Otherwise, the pattern is generated from Number_of_voices and Semitone_step. If Wrap_to_octave is enabled, each step is wrapped into the 0–<12 ST octave.
Lists may be separated by spaces, commas, semicolons, or tabs. The engine accepts up to 32 voices.
Voice rendering
2. ratio = 2^(voice semitones / 12)
3. override sample rate = source rate × ratio
4. resample to Output_sample_rate
5. Scale intensity to the voice's dB SPL target
6. prepend the voice entry delay
7. sum into mono or stereo accumulation buses
Because every Canon voice uses varispeed, transposition changes its active duration. The complete result lasts until the latest entry delay + active voice duration.
Scale intensity is Praat's intensity operation. Start_intensity_dB is the target intensity in dB SPL for voice 1, and Intensity_step_dB changes that target for successive voices. These fields are not simple relative gain offsets.
Spatialisation
- Mono — all voices are summed to one channel.
- Stereo spread — voices are distributed from left to right using constant-power coefficients
cos(angle)andsin(angle). - Alternating L-R — odd/even voices alternate between fixed left-heavy and right-heavy constant-power positions.
Humanize
Humanize_cents adds independent uniform cent jitter to each voice. Humanize_timing_ms adds independent uniform timing jitter to each entry; negative resulting delays are clamped to zero. When Random_seed is greater than zero, the script initializes the random generator with that seed for repeatable humanization.
Presets
Detune presets
Detune presets set interval, balance, and Haas delay. They do not change the selected detune method or the advanced output/analysis settings.
| Preset | Interval | Balance | Haas |
|---|---|---|---|
| Custom | Uses current value | Uses current value | Uses current value |
| Subtle chorus | 0.10 ST | Symmetric split | 0 ms |
| Classic detune | 1.65 ST | Right channel only | 0 ms |
| Wide doubler | 3.00 ST | Symmetric split | +14 ms right |
| Honky-tonk | 0.50 ST | Symmetric split | 0 ms |
| Extreme split | 7.00 ST | Symmetric split | +22 ms right |
Canon presets
Canon presets set only the fields shown below. Values not listed remain at the current/default value. Major and Minor Arpeggio use explicit pitch patterns; the other built-in canons use the step rule.
| Preset | Voices / pattern | Entry delay | Intensity step | Other override |
|---|---|---|---|---|
| Custom | Current settings | Current | Current | None |
| Major arpeggio (fast) | 0 4 7 11 | 0.25 s | −2 dB | Explicit pattern |
| Minor arpeggio | 0 3 7 10 | 0.30 s | −2 dB | Explicit pattern |
| Spooky cluster (slow) | 5 voices · step +1 ST | 1.20 s | −1 dB | Wrap off |
| Octave stacks | 3 voices · step +12 ST | 0.50 s | −2 dB | Wrap off |
| Quartal stack | 4 voices · step +5 ST | 0.40 s | −2 dB | Wrap off |
| Whole-tone cloud | 6 voices · step +2 ST | 0.70 s | −1.5 dB | Wrap off |
| Descending canon | 4 voices · step −3 ST | 0.45 s | −2 dB | Wrap off |
| Shepard spiral | 8 voices · step +7 ST | 0.35 s | −1.5 dB | Wrap on; Stereo spread |
Parameters
Detune main settings
| Parameter | Default before preset | Description |
|---|---|---|
| Stereo_detune_semitones | 1.65 ST | Interval between left and right channels. |
| Detune_balance | Right channel only | Right-only shift or symmetric ±half split. |
| Detune_method | PSOLA | Duration-preserving PSOLA or tape-style varispeed. |
| Haas_delay_ms | 0 ms | Positive delays right; negative delays left; ±200 ms maximum. |
Canon main settings
| Parameter | Default before preset | Description |
|---|---|---|
| Number_of_voices | 4 | Used by the step rule when no explicit list/preset pattern is active. |
| Delay_between_entries | 0.5 s | Nominal spacing between voice entries. |
| Semitone_step | 7 ST | Pitch increment for the step rule; may be negative. |
| Wrap_to_octave | Yes | Wraps step-rule values into 0–<12 ST. |
| Semitone_list | blank | Free list of semitone values; overrides preset pattern, step rule, and voice count. |
| Start_intensity_dB | 70 dB | Praat Scale intensity target for voice 1. |
| Intensity_step_dB | −3 dB | Change in intensity target for each later voice. |
| Canon_spatialisation | Mono | Mono, constant-power Stereo spread, or Alternating L-R. |
Advanced settings
| Parameter | Default | Description |
|---|---|---|
| Output_sample_rate | 44100 Hz | Final rendering rate; validated from 8000 to 384000 Hz. |
| Fade_ms | 10 ms | Raised-cosine fade-in and fade-out on the final output. It is applied only when twice the requested fade is shorter than the result. |
| Peak_ceiling | 0.99 | Final attenuate-only absolute peak ceiling; valid range >0 to 1. |
| Pitch_floor_Hz | 40 Hz | Pitch-analysis floor used for source reporting and Detune PSOLA. |
| Pitch_ceiling_Hz | 1200 Hz | Requested analysis ceiling; automatically clamped to 0.45 × source sample rate. |
| Resample_precision | 50 | Praat resampling precision; valid range 1–1000. |
| Spectrogram_max_Hz | 5000 Hz | Visualization ceiling; clamped to source Nyquist. |
| Humanize_cents | 0 | Canon only: uniform per-voice pitch jitter in cents. |
| Humanize_timing_ms | 0 | Canon only: uniform per-voice entry-time jitter in milliseconds. |
| Random_seed | 20260829 | Canon only: positive values initialize repeatable humanization. |
Output behavior
| Property | Stereo Pitch Detune | Time-Delayed Canon |
|---|---|---|
| Output name | <source>_detune_<preset> | <source>_canon_<preset> |
| Source channels | Folded to mono before processing | Folded to mono before processing |
| Final channels | Always stereo | Mono or stereo according to spatialisation |
| Pitch method | PSOLA or varispeed | Varispeed for every voice |
| Duration | PSOLA preserves active duration; Haas adds delay. Varispeed changes channel durations and shorter channels are tail-padded. | Depends on each voice's transposition and entry delay; output ends at the latest voice end. |
| Sample rate | Final output is rendered at Output_sample_rate. | |
Fade and peak safety
The final Sound receives a raised-cosine fade-in/out when the requested fade fits within the output. The script then measures the absolute extremum. Scale peak is called only if that extremum exceeds Peak_ceiling. Material already below the ceiling is left at its existing level.
Visualization
The v2.1 Picture-window page uses a shared 8-inch layout and contains:
- Original waveform and result waveform on one shared amplitude range, so their levels are directly comparable.
- Original and result spectrograms, using the same requested upper-frequency limit.
- Mode-specific analysis panel:
- Detune — measured F0 curves of the actual left and right output channels. The display range uses 2%/98% F0 quantiles and suppresses isolated out-of-range frames in drawing copies.
- Canon — an entry timeline showing each voice, its nominal semitone value, intensity target, entry delay, and varispeed-dependent bar length.
- Interval structure panel: Detune shows channel interval/ratios, expected mean-F0 values, beat rate and Haas direction; Canon lists interval names and voice entries (up to nine rows in the compact panel).
- Technique panel describing the actual render path, output rate, fade, and peak ceiling.
- Summary strip with the selected preset and principal output statistics.