Polyphonic Improviser — Chunk Shuffle Canon — User Guide

Divides one source into equal chunks, independently shuffles those chunks for 2–4 delayed voices, transforms each voice by varispeed or pitch-preserving time scaling, joins the chunks with fixed 40% equal-power overlap, and pans the resulting mono voices into a newly generated stereo texture.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 2.6 (2026) Technique: Chunk Shuffle Canon Category: Composition License: MIT License Repo: GitHub
Contents:

What this does

Polyphonic Improviser is a source-derived recomposition engine. It creates no oscillator or synthetic note material: every voice is assembled from chunks extracted from the selected Sound.

source → mono fold → N equal chunks → transform each chunk separately for each voice → independent Fisher–Yates shuffle per voice → 40% equal-power overlap-add within each voice → stagger voice entries → voice amplitude → constant-power pan → additive stereo mix → peak normalization → silence trim → final fade-out
Why “canon”? Voice 1 enters first, Voice 2 after one entry delay, Voice 3 after two delays, and Voice 4 after three. Unlike a traditional imitative canon, however, each voice receives a different permutation of the source chunks and can use a different transformation ratio.
Spatial contract: stereo and multichannel input are converted to mono before chunking. The result is a newly generated stereo field created by voice panning. The source's original stereo/multichannel image is not preserved.

Quick start

  1. Select exactly one Sound of at least 1 second.
  2. Run Polyphonic_Improviser.praat.
  3. Choose Custom or one of the six named presets.
  4. For Custom, set chunk count, voice count, transform mode, V1/V2 ratios and entry timing in the main form.
  5. For Custom only, the second dialog exposes V3/V4 ratios plus all voice amplitudes and pans.
  6. Run the script. The result is named <source>_poly_improv_v2.
Named presets skip the Voice Details dialog because they explicitly replace the voice ratios, amplitudes and pans. The dialog appears only for Custom.

Source preparation & chunking

The selected Sound is copied to a mono working source. Multichannel input uses Praat Convert to mono; mono input is simply copied. The working source is then shifted to start at time 0 so non-zero source xmin values do not affect chunk coordinates.

The source is divided into equal temporal chunks:

chunkDuration = sourceDuration / Number_of_chunks chunk c: start = (c - 1) × chunkDuration end = c × chunkDuration

Extraction is rectangular. There is no transient detection, beat detection, TextGrid segmentation, or variable chunk sizing.

Number_of_chunks is clamped to 2–200.

Shuffle engine

Each voice begins with the source-order index list:

[1, 2, 3, ..., N]

Before shuffling voice v, the script consumes (v-1) × N random integer draws. It then performs a standard Fisher–Yates permutation from the end of the array toward the beginning.

for i = N ... 2: j = randomInteger(1, i) swap(index[i], index[j])

The additional random draws make the voices consume different regions of Praat's random stream. The actual permutation is stored and later reused by the visualization, so the shuffle map always shows the ordering that produced the audio.

The shuffles are not reproducible from the form settings alone. There is no random-seed control and the script does not initialize a fixed seed. Re-running the same settings can therefore produce different chunk orders.

Transformation modes

Every source chunk is transformed separately for every active voice using that voice's speed ratio.

Tape speed — pitch and time together

Override sampling frequency: sourceSR × voiceRatio then: Resample back to sourceSR, precision 50

This is true varispeed behavior:

Lengthen — time only

lengthenFactor = 1 / voiceRatio Praat: Lengthen (overlap-add): 75, 600, lengthenFactor

This is Praat's pitch-preserving overlap-add time scaling. A ratio above 1 requests a shorter result; a ratio below 1 requests a longer result while attempting to preserve pitch.

Voice ratios are clamped to 0.05–8.0. In Lengthen mode the reciprocal factor is then clamped internally to 0.1–8.0. With the current ratio clamp, the practically reachable factor is approximately 0.125–8.0.

Lengthen quality depends on the suitability of the material for Praat's periodicity-based overlap-add algorithm. It is not a phase-vocoder or granular time-stretch engine.

Fixed 40% equal-power legato overlap

The current engine does not use the form's Crossfade_ms value for chunk joins. Since v2.3, every transformed chunk uses a fixed overlap equal to 40% of that voice's transformed chunk duration.

overlap = 0.40 × transformedChunkDuration hop = transformedChunkDuration - overlap = 0.60 × transformedChunkDuration

Before placement, every chunk receives sinusoidal edge windows over its first and last 40%:

fade-in: sin(progress × π/2) fade-out: sin(reverseProgress × π/2)

Adjacent chunks are placed one 60% hop apart, so the outgoing and incoming windows coincide. Their squared gains sum to 1 through the overlap, producing the intended equal-power legato join.

Crossfade_ms is currently a legacy compatibility field. It remains in the public form and presets so older runScript calls keep the same signature, but changing it does not alter the v2.6 audio join. The actual overlap is always 40%.

The first chunk also fades in from zero and the final chunk fades out toward zero as part of the same windowing scheme.

Voice entries & timing

Voice entry times are:

V1 = 0 V2 = 1 × entryDelay V3 = 2 × entryDelay V4 = 3 × entryDelay

Manual

When Quantize_entries is off, Entry_delay_s is used directly in seconds.

BPM-referenced

When Quantize_entries is on:

beatDuration = 60 / Tempo_bpm entryDelay = beatDuration × noteBeats
Note valueBeats used by the script
Whole4
Half2
Quarter1
Eighth0.5
Dotted whole6
Dotted half3
Dotted quarter1.5
2 bars8
4 bars16
This is entry-time arithmetic, not tempo or beat detection. The script never analyzes the source for BPM, meter, beats or downbeats.

Voice amplitude & stereo pan

After a voice's shuffled mono sequence has been assembled and padded to the common master duration, its amplitude multiplier is applied.

Panning then uses a constant-power law:

angle = (pan + 1) / 2 × π/2 leftGain = cos(angle) rightGain = sin(angle)
PanResult
−1full left, zero right
0approximately 0.707 left and 0.707 right
+1zero left, full right

All left voice signals are summed together, all right voice signals are summed together, and those two buses are combined into the final stereo Sound.

Voice amplitudes are clamped at a minimum of 0; pans are clamped to −1…+1.

Presets

Named presets overwrite chunk count, the legacy Crossfade_ms field, voice count, transform mode, voice ratios, amplitudes, pans, and entry-timing behavior. They also skip the Custom Voice Details dialog.

PresetChunksVoicesModeRatiosEntry
Slow Canon83Tape speed1.000, 1.059, 0.5004 bars @ 72 BPM = 13.333 s
Dense Cluster244Tape speed1.000, 1.059, 1.122, 0.944quarter @ 120 BPM = 0.500 s
Spectral Drift63Lengthen1.000, 0.850, 1.200manual 6.0 s
Rhythmic Echo164Tape speed1.000, 1.498, 0.500, 0.749half @ 100 BPM = 1.200 s
Mirror Scatter204Tape speed1.000, 1.000, 1.330, 0.750quarter @ 90 BPM = 0.667 s
Microtonal Haze104Lengthen1.000, 1.025, 0.975, 1.050manual 1.5 s
The Slow Canon preset's 4 bars setting is implemented literally as 16 quarter-note beats. At 72 BPM that is approximately 13.33 seconds between voice entries, not 8 seconds.

Draw_visualization and Play_output are not overwritten by the presets.

Parameters & effective limits

Main form

ParameterDefaultBehavior
PresetCustomCustom plus six named strategies.
Number_of_chunks12Clamped to 2–200.
Crossfade_ms30Legacy API field; clamped 0–500 but does not control current joins.
Number_of_voices3Clamped to 2–4; clamp is reported.
Transform_modeTape speedVarispeed or pitch-preserving Lengthen.
V1_speed_ratio1.00Voice 1 transformation ratio.
V2_speed_ratio1.059Voice 2 transformation ratio.
Quantize_entriesOffChoose BPM/note arithmetic instead of manual seconds.
Tempo_bpm120Reference tempo used for entry arithmetic and visualization beat grid.
Note_valueQuarterBeat multiplier when quantized entry mode is active.
Entry_delay_s3.0Manual inter-voice entry delay.
Draw_visualizationOnDraw the 8×8 process visualization.
Play_outputOnPlay final result.

Custom Voice Details dialog

ParameterDefault
V3_speed_ratio0.50
V4_speed_ratio1.50
V1_amplitude1.00
V2_amplitude0.85
V3_amplitude0.75
V4_amplitude0.65
V1_pan−0.35
V2_pan+0.40
V3_pan−0.75
V4_pan+0.75

Speed ratios are clamped to 0.05–8.0; amplitudes to ≥0; pans to −1…+1.

Duration, normalization, silence trim & final fade

Measured voice body duration

All transformed chunks within one voice have the same duration. With a 40% overlap:

voiceBody = transformedChunkDuration + (N - 1) × 0.60 × transformedChunkDuration

The pre-mix master duration is based on the latest actual voice end:

voiceEnd[v] = (v - 1) × entryDelay + voiceBody[v] masterDuration = max(voiceEnd[v]) + 0.5 s safety pad

This v2.6 calculation prevents slower voices from being truncated merely because Voice 1 is shorter.

Peak normalization

After stereo mixing, the script measures the Sinc70 absolute peak. For any non-silent result it calls:

Scale peak: 0.95

This is target peak normalization, not attenuation-only safety limiting. A quiet non-zero mix can be amplified to a 0.95 Sinc70 peak.

Automatic silence trim

The normalized stereo result is folded to mono for a 50 ms amplitude scan. A block counts as active if:

max > 0.005 or min < -0.005

The script keeps approximately 50 ms of margin around the first and last detected active block. It performs the extraction only when more than 100 ms of silence would actually be removed from the beginning or end.

If the entire result remains below the threshold, trimming is skipped.

Final fade-out

After trimming, a raised-cosine fade to exact silence is applied over the last:

min(2.0 s, 50% of final duration)

There is no second peak-normalization pass after trim or fade. The delivered output can therefore peak below 0.95 if the trimmed-away region contained the previous peak or if the final fade attenuates it.

Visualization

The current v2.6 Picture view follows the actual recomposition process:

  1. Chunk shuffle map: top source row shows chunks 1…N; each voice row shows the stored permutation that was actually rendered.
  2. Active span per voice: uses each voice's measured transformed OLA body duration, corrected for any leading trim offset.
  3. Left waveform: final channel 1.
  4. Right waveform: final channel 2.
  5. Summary strip: preset, source, voices, chunks, transform, fixed 40% overlap, voice ratios/pans, entry timing, output name and final duration.

Waveform scaling

Left and Right share the same amplitude range, calculated from the larger of the two channel peaks. Their displayed levels can therefore be compared directly.

Beat grid

The light dotted BPM grid is drawn even when Quantize_entries is off. In Manual mode it is only a reference grid derived from Tempo_bpm; it does not mean that the actual entry delay was quantized.

Voice-entry markers

The colored dotted entry lines use the actual delayed entries minus any leading silence removed by auto-trim, so the markers remain aligned to the delivered output.

Output behavior

The final output combines three distinct operations that should not be conflated: chunk permutation changes order, voice transformation changes chunk time/pitch behavior, and entry delay + pan creates the polyphonic stereo layout.