Perceptual Fugue — User Guide

A fugue-inspired stereo construction engine that turns one source into recurring, transposed, reversed, augmented, fragmented, and overlapping entries. Each voice keeps a fixed spatial signature while its musical material changes with the section.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 2.6.3 (2026) Category: Spatial & Immersive Audio
Contents:

What this does

Perceptual Fugue builds a fixed five-section timeline from one selected Sound. It is best described as a fugue-inspired polyphonic audio construction, not as an automatic implementation of strict classical fugue rules. The source acts as the subject; derived materials include answer transpositions, an octave-lower subject, counter-subject material, a half-length head motif, optional retrograde, optional 2× augmentation, and stretto entries.

After the musical timelines are assembled, each voice receives a fixed stereo signature built from inter-channel delay, level asymmetry, optional far-side low-pass filtering, reverb/reflections, and a shared mid-band identity anchor. The final output is stereo.

Input: exactly one mono or stereo Sound, at least 0.1 s long. Stereo input is converted to mono before any subject materials are created. If the Sound begins at a non-zero time, the working mono copy is re-extracted so its timeline begins at 0.
Spatial interpretation: the configured delays are compositional stereo delays, not a physical head model. Values around 2–4 ms are much larger than natural geometric interaural delays, so they function more like precedence/micro-echo cues than natural ITD simulation. The script does not use HRTFs or individualized binaural filters.

Quick start

  1. Select exactly one mono or stereo Sound in Praat.
  2. Run Perceptual_Fugue.praat.
  3. Start with Classical Fugue, the default preset.
  4. Keep Varispeed + PSOLA compensation unless you specifically want the PitchTier method or the legacy varispeed-only behavior.
  5. Set an analysis range that brackets the source pitch when using PSOLA-based methods.
  6. Enable Draw_visualization to see the actual entry score, output waveforms, spectrogram, and summary.
  7. Click OK. The output is named source_perceptual_fugue, peak-normalized to 0.99, and played automatically.
Duration: the musical form ends at [9 + (voices−1)×stretto_compression] × source_duration. A separate absolute tail allocation is then added so fixed reverb/reflection delays are not cut off. The tail setting does not set reverb time.

Form & presets

Current form

ParameterDefaultActual behavior
PresetClassical FugueCustom or one of four fixed strategies. Presets override voices, interval, stretto, retrograde, augmentation, base ITD/ILD, and shadow cutoff.
Answer_intervalFifth upFifth = 2^(7/12), fourth = 2^(5/12), octave = 2, tritone = 2^(6/12).
Number_of_voices3Clamped to 3–4.
Stretto_compression0.75Clamped to 0.2–1.0. Lower values place stretto entries closer together.
Include_retrogradeyesIf off, the counter-subject becomes an amplitude-modulated subject and the middle-entry “retrograde” becomes an unchanged subject copy.
Include_augmentationyesIf on, the middle V3 subject and the pedal drone use 2× duration Manipulation resynthesis. If off, ordinary-duration copies are used.
Transposition_methodVarispeed + PSOLA compensationThree distinct algorithms; see below.
Pitch_floor / Pitch_ceiling50 / 800 HzMust satisfy floor < ceiling. Used by Manipulation-based pitch/duration processing.
Exposition_ITD_ms3.0 msBase delay used by V1/V2; V3 uses 0.33× this value and V4 uses 0.
Exposition_ILD_factor4.0Base V1/V2 level ratio. Values below 1 are inverted automatically so V1 remains left-biased and V2 right-biased.
Shadow_cutoff_Hz500 HzFar-side Hann low-pass cutoff for V1/V2. Capped automatically to 0.95×Nyquist.
Ild_lawConstant powerConstant-energy direct-path ratio or legacy near-side boost.
Tail_allocation_seconds1.5 sTimeline room after the last musical entry; raised automatically if shorter than the longest internal delay/reverb requirement.
Speed_modeBalancedSets varispeed resample precision: 50 / 20 / 10 for Full / Balanced / Fast.
Draw_visualizationyesDraws the score-driven spatial overview after rendering.

Preset strategies

PresetVoices / intervalStrettoSpatial baseRetro / Aug
Classical Fugue3 / fifth0.753.0 ms, ILD 4.0, shadow 500 Hzon / on
Spectral Fugue4 / tritone0.504.0 ms, ILD 5.0, shadow 400 Hzon / on
Chamber Fugue3 / fifth0.852.0 ms, ILD 2.5, shadow 800 Hzon / off
Stretto Study4 / fifth0.403.5 ms, ILD 4.5, shadow 450 Hzon / on

Presets do not override the transposition method, pitch range, ILD law, tail allocation, speed mode, or visualization toggle.

Transposition methods

1. Varispeed + PSOLA compensation

First moves pitch and duration together by overriding sampling frequency and resampling. Then a Manipulation/DurationTier stretch compensates the duration so the whole subject returns to its intended length. The pitch-analysis range is scaled by the transposition ratio, with internal safety limits. Because the compensation still depends on pitch tracking, polyphonic, noisy, or percussive material can produce artifacts.

2. PitchTier

Creates a Manipulation from the source, multiplies PitchTier frequencies by the requested ratio, replaces the tier, and resynthesizes. Duration is preserved by construction. This is most appropriate when a reliable monophonic/voiced pitch track exists.

3. Varispeed only — legacy

Keeps the v2.3 sound intentionally. The resampled result is forced back to the original duration: upward transpositions are followed by padded silence, while downward transpositions are truncated. It therefore does not preserve the complete subject.

Materials & formal timeline

Let S be the source duration. The script creates a subject and several derived materials before building 3 or 4 mono voice timelines.

MaterialConstructionNominal role
subjectSource mono copy with 5 ms raised-cosine edge fadesDux / final statement
answerSubject transposed by selected answer ratioComes
subjectLowSubject ×0.5 pitchLower-register entry
answerLowSubject ×(answer ratio ×0.5); only for four voicesV4 lower answer
counterSubjectRetrograde subject when retrograde is enabled; otherwise amplitude-modulated subjectCounter-subject material
headMotifFirst 50% of subjectEpisode fragment
headMotifT1 / T2Head motif transposed by answerRatio and answerRatio²Episode fragments
augSubject2× duration subject when augmentation is enabled; otherwise subject copyMiddle-entry V3
pedalDrone2× duration subjectLow when augmentation is enabled; otherwise ordinary subjectLow copyPedal

Five sections

Exposition end = 3S Episode end = 4S Middle entries end = 6S Stretto duration = (voices − 1) × compression × S + S Stretto end = 6S + stretto duration Last musical end = stretto end + 2S Output duration = last musical end + tail allocation

Actual placement logic

The construction is richer than a simple “S → A → S” summary. The current score visualization is generated from the same placement log used while these entries are built.

I. Exposition: V1 subject at 0, V1 counter-subject at S, then filtered support at 2S; V2 answer at S and counter-subject at 2S; V3 subjectLow at 2S; optional V4 filtered support at 2S.

II. Episode: four half-subject fragments move between V1 and V2 at quarter-S offsets, while V3 carries a quiet low-passed subjectLow support.

III. Middle entries: V1 retroSubject at 4S and counter-subject one S later; V2 receives the answer transposed by the answer ratio again; V3 receives augSubject; optional V4 receives answerLow.

IV. Stretto: subject-family entries enter at compression×S intervals in V1–V4, with additional counter-subject entries in V1 and V2.

V. Pedal/cadence: V3 pedalDrone begins at the stretto end; V2 adds filtered answer support; V1 gives a final subject one S later; optional V4 adds filtered answerLow support. A raised-cosine cadence fade occupies the last 0.5S of musical time and reaches zero.

Voice identity: the spatial signature is fixed per voice, but pitch/register is not. A voice can carry different derived materials in different sections. “V1 = original pitch” and similar labels describe its principal role, not every event assigned to that voice.

Spatial processing

Before spatialization, each voice timeline is multiplied by 1 / number_of_voices. Each mono voice is then copied to L/R and processed with its own fixed parameter set.

Direct-path ILD law

For requested near/far ratio F ≥ 1: near = √2 × F / √(1 + F²) far = √2 / √(1 + F²) Therefore: near / far = F near² + far² = 2

This keeps the direct stereo-path energy constant relative to a duplicated mono center while changing the L/R ratio. It does not calibrate the energy of the complete spatial chain: filtering, wet signal, reflections, identity anchor, and V3 polarity inversion still alter the final result. The alternative legacy law simply boosts the declared near side by the factor.

Per-voice signatures

VoiceDirect cueShadowReverb / reflections / anchorSpecial behavior
V1 DuxRight delayed by base ITD; left-biased ILDRight low-pass at Shadow_cutoffRT 0.35 s; reverb coefficient 0.12; reflection gain 0.20; anchor 0.08Strongly left-biased, but not hard-panned
V2 ComesLeft delayed by base ITD; right-biased mirror ILDLeft low-passRT 0.35 s; reverb coefficient 0.12; reflection gain 0.20; anchor 0.08Strongly right-biased, but not hard-panned
V3 ThirdRight delayed by 0.33×base ITD; ILD factor 1.3NoneRT 0.50 s; reverb coefficient 0.35; reflection gain 0.15; anchor 0.10Right channel polarity inverted deliberately; poor mono compatibility is expected
V4 FourthNo ITD; unity ILDNoneRT 0.80 s; reverb coefficient 0.60; reflection gain 0.10; anchor 0.12Centered direct/shared-wet field with side-assigned reflections; not a physically diffuse-field model

Processing order inside each voice

  1. ITD-style delay: delay the designated far/lagging channel.
  2. ILD: constant-energy direct-path law or legacy near-side boost.
  3. Spectral shadow: V1/V2 far side receives Praat Filter (pass Hann band) from 0 Hz to the cutoff with 100 Hz smoothing.
  4. Polarity: V3 right channel is multiplied by −1.
  5. Reverb bus: taps at 0.25×RT, 0.5×RT (×0.65), and 1.0×RT (×0.35), then low-pass-smoothed. The form does not expose these fixed per-voice values.
  6. Early/late reflections: early taps at 0.2×RT and 0.4×RT; a late filtered tap at 1.0×RT. They are assigned to opposite/same sides according to voice delay direction.
  7. Identity anchor: a 500–3000 Hz Hann-band copy of the mono voice is added equally to both channels.
Reverb coefficient is not a literal wet percentage. The wet bus contains several additive taps whose raw weights sum to 2.0 before filtering. The code also scales the direct L/R path by 1 − coefficient, so the coefficient controls the dry/wet construction but should not be interpreted as a measured percentage of output energy.

Mixing, normalization & mono report

The stereo voices are summed channel-by-channel with additive formulas. The result is renamed source_perceptual_fugue and then processed with Scale peak: 0.99.

Peak scaling is full normalization. It can attenuate a hot mix or amplify a quiet one. The older Section_intensity_dB control was removed because an intensity scaling immediately followed by peak normalization was contradictory.

Mono-fold diagnostic

Because V3 deliberately inverts its right-channel polarity, the script measures what happens when the final stereo output is folded to mono.

M = (L + R) / 2 Eref = (EL + ER) / 4 mono-fold ratio = E(M) / Eref Anchors used by the report: 200% = identical equal-energy channels in phase 100% = uncorrelated reference 0% = equal-energy anti-phase cancellation Cross term: C = [4 E(M) − EL − ER] / 2 Zero-lag normalized correlation: ρ = C / √(EL ER)

The 0–200% figure is not “energy retained.” It is referenced to the energy expected from an uncorrelated L/R pair. The normalized correlation is reported separately on the standard −1…+1 scale. If either channel is effectively silent, correlation is reported as undefined.

The script uses its mono-fold ratio as a practical warning heuristic: below 60% it reports substantial cancellation; 60–95% reports some cancellation; at or above 95% it reports little cancellation on that material. These are tool-specific report thresholds, not a general broadcast or mastering standard.

Visualization

The current v2.6.3 visualization is a process-oriented overview. It does not contain the old L−R-difference panel described by earlier documentation.

1. Entry score

One row is drawn per active voice. Crucially, the blocks are no longer maintained as a separate hand-written diagram: every timeline placement is logged at the same time the fragment is inserted, and the score reads that placement log. Block lengths therefore use the actual duration of each placed fragment, including changes caused by Retrograde/Augmentation settings.

S subject-family A answer-family CS counter-subject R retrograde Aug augmented frag head-motif fragment pad filtered support Ped pedal

Vertical section lines mark Exposition, Episode, Middle entries, Stretto, and Pedal. The cadence-fade region and the post-music tail allocation are shaded separately.

2. Left and right waveforms

Separate waveform panels use a shared amplitude scale so channel-level differences remain visually comparable. Section boundaries are overlaid.

3. Spectrogram

The spectrogram is calculated from the already-extracted left channel, because Praat’s spectrogram operation used here is mono-only. It covers 0–5000 Hz and carries the same formal boundaries.

4. Summary

The final compact panel reports preset, source, voice count, answer interval, output duration, and the five-part formal sequence.

What the visualization does not claim: score blocks show scheduled source-material placements, not an acoustic source-separation analysis of the final reverberant mix. The waveform and spectrogram show the rendered output, while the score shows the construction that produced it.

Limits & interpretation

Applications

Compositional transformation

Turn a short pitched gesture into a longer form whose repeated identity is distributed across register, time, and stereo signatures.

Spatial counterpoint studies

Compare how fixed left/right/center-reverberant signatures interact with recurring musical materials and stretto density.

Analysis and teaching

Use the score-driven visualization to relate scheduled entries to the rendered stereo waveform and spectrogram.