Pitch Processor — User Guide

Two complementary pitch-based processors: a stereo detune engine with PSOLA or varispeed transposition, and a time-delayed canon built from independently transposed voices.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 2.1 (2026) License: MIT License Repo: https://github.com/ShaiCohen-ops/Praat-plugin_AudioTools
Contents:

What this does

Pitch Processor contains two independent processing modes:

Source-channel behavior: the selected Sound is converted to a mono working copy before either mode is rendered. A stereo or multichannel source therefore does not retain its original spatial image. Detune always constructs a new stereo result; Canon constructs mono or stereo according to Canon_spatialisation.

Both modes end with an optional raised-cosine fade and an attenuate-only peak ceiling. Quiet results are not normalized upward.

Quick start

  1. Select exactly one Sound object.
  2. Run Pitch_Processor.praat.
  3. Choose Mode: Stereo Pitch Detune or Time-Delayed Canon.
  4. Choose the preset for the selected mode.
  5. Choose Settings:
    • Run the preset as it is — render immediately with the preset/default values.
    • Edit main settings — open only the controls relevant to the selected mode.
    • Edit main and advanced settings — open the mode controls, then output/analysis controls.
  6. Choose whether to draw the visualization and play the result.
  7. Click OK.

Interface & settings dialogs

The main form is intentionally compact. It contains only the mode, the mode-specific preset menus, the settings switch, and the draw/play toggles. Processing parameters live in optional dialogs so that Detune and Canon do not expose irrelevant controls.

Main form

FieldDefaultMeaning
ModeStereo Pitch DetuneSelects the processing engine.
Detune_presetClassic detuneUsed when Mode = Stereo Pitch Detune.
Canon_presetCustomUsed when Mode = Time-Delayed Canon.
SettingsRun the preset as it isControls whether the optional parameter dialogs appear.
Draw_visualizationYesDraws the 8-inch analysis page in the Picture window.
Play_after_processingYesPlays the completed result.
Preset patterns: Major Arpeggio and Minor Arpeggio contain explicit pitch lists (0 4 7 11 and 0 3 7 10). When one of these presets is selected, that stored pattern takes precedence over the step rule unless you enter your own Semitone_list. A user Semitone_list always has highest priority and also determines the number of voices.

Mode 1 — Stereo Pitch Detune

The Detune engine creates two transposed copies of the mono working source and combines them as left and right channels. Stereo_detune_semitones specifies the interval between the two channels.

Balance

Detune_balanceLeft shiftRight shift
Right channel only0 ST+Detune ST
Symmetric split−Detune/2+Detune/2

PSOLA (duration preserving)

For a non-zero channel shift, Praat creates a Manipulation object, extracts its PitchTier, multiplies all defined F0 values by 2^(semitones/12), replaces the tier, and resynthesizes by overlap-add. The result is then resampled to Output_sample_rate. The pitch change preserves the source duration before any Haas delay is added.

Automatic fallback: if the source is too short for the requested pitch analysis or contains no voiced frames, PSOLA is not attempted. The script switches the Detune render to Varispeed and reports the fallback in the Info window.

Varispeed (tape transposition)

Varispeed overrides the working copy's sampling frequency by the pitch ratio, then resamples it to the requested output rate. Pitch and duration therefore change together:

ratio = 2^(semitones / 12)
active duration ≈ source duration / ratio

The two channels are padded to the same final duration before stereo combination. With unequal varispeed shifts, this padding does not restore time alignment; it only makes the channel lengths equal.

Haas delay

Haas_delay_ms prepends silence to one channel after pitch processing. Positive values delay the right channel; negative values delay the left. The absolute value is limited to 200 ms. A Haas delay therefore increases the final output duration.

The Info window also reports the nominal beat rate at the measured source mean F0 when a usable mean F0 exists.

Mode 2 — Time-Delayed Canon

The Canon engine creates one mono varispeed copy per voice. Each voice has a semitone value, optional cent jitter, entry delay, intensity target, and pan coefficients.

Pitch-pattern priority

  1. If Semitone_list is non-empty, it becomes the complete pitch pattern and its length becomes the voice count.
  2. Otherwise, if the selected preset contains an explicit pattern, that preset pattern is used.
  3. Otherwise, the pattern is generated from Number_of_voices and Semitone_step. If Wrap_to_octave is enabled, each step is wrapped into the 0–<12 ST octave.

Lists may be separated by spaces, commas, semicolons, or tabs. The engine accepts up to 32 voices.

Voice rendering

1. voice semitones = pattern value + optional humanize cents / 100
2. ratio = 2^(voice semitones / 12)
3. override sample rate = source rate × ratio
4. resample to Output_sample_rate
5. Scale intensity to the voice's dB SPL target
6. prepend the voice entry delay
7. sum into mono or stereo accumulation buses

Because every Canon voice uses varispeed, transposition changes its active duration. The complete result lasts until the latest entry delay + active voice duration.

Intensity semantics: Scale intensity is Praat's intensity operation. Start_intensity_dB is the target intensity in dB SPL for voice 1, and Intensity_step_dB changes that target for successive voices. These fields are not simple relative gain offsets.

Spatialisation

Humanize

Humanize_cents adds independent uniform cent jitter to each voice. Humanize_timing_ms adds independent uniform timing jitter to each entry; negative resulting delays are clamped to zero. When Random_seed is greater than zero, the script initializes the random generator with that seed for repeatable humanization.

Presets

Detune presets

Detune presets set interval, balance, and Haas delay. They do not change the selected detune method or the advanced output/analysis settings.

PresetIntervalBalanceHaas
CustomUses current valueUses current valueUses current value
Subtle chorus0.10 STSymmetric split0 ms
Classic detune1.65 STRight channel only0 ms
Wide doubler3.00 STSymmetric split+14 ms right
Honky-tonk0.50 STSymmetric split0 ms
Extreme split7.00 STSymmetric split+22 ms right

Canon presets

Canon presets set only the fields shown below. Values not listed remain at the current/default value. Major and Minor Arpeggio use explicit pitch patterns; the other built-in canons use the step rule.

PresetVoices / patternEntry delayIntensity stepOther override
CustomCurrent settingsCurrentCurrentNone
Major arpeggio (fast)0 4 7 110.25 s−2 dBExplicit pattern
Minor arpeggio0 3 7 100.30 s−2 dBExplicit pattern
Spooky cluster (slow)5 voices · step +1 ST1.20 s−1 dBWrap off
Octave stacks3 voices · step +12 ST0.50 s−2 dBWrap off
Quartal stack4 voices · step +5 ST0.40 s−2 dBWrap off
Whole-tone cloud6 voices · step +2 ST0.70 s−1.5 dBWrap off
Descending canon4 voices · step −3 ST0.45 s−2 dBWrap off
Shepard spiral8 voices · step +7 ST0.35 s−1.5 dBWrap on; Stereo spread

Parameters

Detune main settings

ParameterDefault before presetDescription
Stereo_detune_semitones1.65 STInterval between left and right channels.
Detune_balanceRight channel onlyRight-only shift or symmetric ±half split.
Detune_methodPSOLADuration-preserving PSOLA or tape-style varispeed.
Haas_delay_ms0 msPositive delays right; negative delays left; ±200 ms maximum.

Canon main settings

ParameterDefault before presetDescription
Number_of_voices4Used by the step rule when no explicit list/preset pattern is active.
Delay_between_entries0.5 sNominal spacing between voice entries.
Semitone_step7 STPitch increment for the step rule; may be negative.
Wrap_to_octaveYesWraps step-rule values into 0–<12 ST.
Semitone_listblankFree list of semitone values; overrides preset pattern, step rule, and voice count.
Start_intensity_dB70 dBPraat Scale intensity target for voice 1.
Intensity_step_dB−3 dBChange in intensity target for each later voice.
Canon_spatialisationMonoMono, constant-power Stereo spread, or Alternating L-R.

Advanced settings

ParameterDefaultDescription
Output_sample_rate44100 HzFinal rendering rate; validated from 8000 to 384000 Hz.
Fade_ms10 msRaised-cosine fade-in and fade-out on the final output. It is applied only when twice the requested fade is shorter than the result.
Peak_ceiling0.99Final attenuate-only absolute peak ceiling; valid range >0 to 1.
Pitch_floor_Hz40 HzPitch-analysis floor used for source reporting and Detune PSOLA.
Pitch_ceiling_Hz1200 HzRequested analysis ceiling; automatically clamped to 0.45 × source sample rate.
Resample_precision50Praat resampling precision; valid range 1–1000.
Spectrogram_max_Hz5000 HzVisualization ceiling; clamped to source Nyquist.
Humanize_cents0Canon only: uniform per-voice pitch jitter in cents.
Humanize_timing_ms0Canon only: uniform per-voice entry-time jitter in milliseconds.
Random_seed20260829Canon only: positive values initialize repeatable humanization.

Output behavior

PropertyStereo Pitch DetuneTime-Delayed Canon
Output name<source>_detune_<preset><source>_canon_<preset>
Source channelsFolded to mono before processingFolded to mono before processing
Final channelsAlways stereoMono or stereo according to spatialisation
Pitch methodPSOLA or varispeedVarispeed for every voice
DurationPSOLA preserves active duration; Haas adds delay. Varispeed changes channel durations and shorter channels are tail-padded.Depends on each voice's transposition and entry delay; output ends at the latest voice end.
Sample rateFinal output is rendered at Output_sample_rate.

Fade and peak safety

The final Sound receives a raised-cosine fade-in/out when the requested fade fits within the output. The script then measures the absolute extremum. Scale peak is called only if that extremum exceeds Peak_ceiling. Material already below the ceiling is left at its existing level.

Visualization

The v2.1 Picture-window page uses a shared 8-inch layout and contains:

Canon visualization and humanize: the timeline's colour and nominal semitone label come from the un-humanized pitch pattern. Humanize cents and timing are reported in the Info window per voice; timing jitter is reflected in the actual entry position because the timeline uses the rendered voice delays.