Polyphonic Improviser — Chunk Shuffle Canon — User Guide
Divides one source into equal chunks, independently shuffles those chunks for 2–4 delayed voices, transforms each voice by varispeed or pitch-preserving time scaling, joins the chunks with fixed 40% equal-power overlap, and pans the resulting mono voices into a newly generated stereo texture.
What this does
Polyphonic Improviser is a source-derived recomposition engine. It creates no oscillator or synthetic note material: every voice is assembled from chunks extracted from the selected Sound.
Quick start
- Select exactly one Sound of at least 1 second.
- Run
Polyphonic_Improviser.praat. - Choose Custom or one of the six named presets.
- For Custom, set chunk count, voice count, transform mode, V1/V2 ratios and entry timing in the main form.
- For Custom only, the second dialog exposes V3/V4 ratios plus all voice amplitudes and pans.
- Run the script. The result is named
<source>_poly_improv_v2.
Source preparation & chunking
The selected Sound is copied to a mono working source. Multichannel input uses Praat Convert to mono; mono input is simply copied. The working source is then shifted to start at time 0 so non-zero source xmin values do not affect chunk coordinates.
The source is divided into equal temporal chunks:
Extraction is rectangular. There is no transient detection, beat detection, TextGrid segmentation, or variable chunk sizing.
Number_of_chunks is clamped to 2–200.
Shuffle engine
Each voice begins with the source-order index list:
Before shuffling voice v, the script consumes (v-1) × N random integer draws. It then performs a standard Fisher–Yates permutation from the end of the array toward the beginning.
The additional random draws make the voices consume different regions of Praat's random stream. The actual permutation is stored and later reused by the visualization, so the shuffle map always shows the ordering that produced the audio.
Transformation modes
Every source chunk is transformed separately for every active voice using that voice's speed ratio.
Tape speed — pitch and time together
This is true varispeed behavior:
- ratio > 1 → shorter and higher;
- ratio = 1 → original pitch/duration;
- ratio < 1 → longer and lower.
Lengthen — time only
This is Praat's pitch-preserving overlap-add time scaling. A ratio above 1 requests a shorter result; a ratio below 1 requests a longer result while attempting to preserve pitch.
Voice ratios are clamped to 0.05–8.0. In Lengthen mode the reciprocal factor is then clamped internally to 0.1–8.0. With the current ratio clamp, the practically reachable factor is approximately 0.125–8.0.
Fixed 40% equal-power legato overlap
The current engine does not use the form's Crossfade_ms value for chunk joins. Since v2.3, every transformed chunk uses a fixed overlap equal to 40% of that voice's transformed chunk duration.
Before placement, every chunk receives sinusoidal edge windows over its first and last 40%:
Adjacent chunks are placed one 60% hop apart, so the outgoing and incoming windows coincide. Their squared gains sum to 1 through the overlap, producing the intended equal-power legato join.
runScript calls keep the same signature, but changing it does not alter the v2.6 audio join. The actual overlap is always 40%.
The first chunk also fades in from zero and the final chunk fades out toward zero as part of the same windowing scheme.
Voice entries & timing
Voice entry times are:
Manual
When Quantize_entries is off, Entry_delay_s is used directly in seconds.
BPM-referenced
When Quantize_entries is on:
| Note value | Beats used by the script |
|---|---|
| Whole | 4 |
| Half | 2 |
| Quarter | 1 |
| Eighth | 0.5 |
| Dotted whole | 6 |
| Dotted half | 3 |
| Dotted quarter | 1.5 |
| 2 bars | 8 |
| 4 bars | 16 |
Voice amplitude & stereo pan
After a voice's shuffled mono sequence has been assembled and padded to the common master duration, its amplitude multiplier is applied.
Panning then uses a constant-power law:
| Pan | Result |
|---|---|
| −1 | full left, zero right |
| 0 | approximately 0.707 left and 0.707 right |
| +1 | zero left, full right |
All left voice signals are summed together, all right voice signals are summed together, and those two buses are combined into the final stereo Sound.
Voice amplitudes are clamped at a minimum of 0; pans are clamped to −1…+1.
Presets
Named presets overwrite chunk count, the legacy Crossfade_ms field, voice count, transform mode, voice ratios, amplitudes, pans, and entry-timing behavior. They also skip the Custom Voice Details dialog.
| Preset | Chunks | Voices | Mode | Ratios | Entry |
|---|---|---|---|---|---|
| Slow Canon | 8 | 3 | Tape speed | 1.000, 1.059, 0.500 | 4 bars @ 72 BPM = 13.333 s |
| Dense Cluster | 24 | 4 | Tape speed | 1.000, 1.059, 1.122, 0.944 | quarter @ 120 BPM = 0.500 s |
| Spectral Drift | 6 | 3 | Lengthen | 1.000, 0.850, 1.200 | manual 6.0 s |
| Rhythmic Echo | 16 | 4 | Tape speed | 1.000, 1.498, 0.500, 0.749 | half @ 100 BPM = 1.200 s |
| Mirror Scatter | 20 | 4 | Tape speed | 1.000, 1.000, 1.330, 0.750 | quarter @ 90 BPM = 0.667 s |
| Microtonal Haze | 10 | 4 | Lengthen | 1.000, 1.025, 0.975, 1.050 | manual 1.5 s |
Draw_visualization and Play_output are not overwritten by the presets.
Parameters & effective limits
Main form
| Parameter | Default | Behavior |
|---|---|---|
| Preset | Custom | Custom plus six named strategies. |
| Number_of_chunks | 12 | Clamped to 2–200. |
| Crossfade_ms | 30 | Legacy API field; clamped 0–500 but does not control current joins. |
| Number_of_voices | 3 | Clamped to 2–4; clamp is reported. |
| Transform_mode | Tape speed | Varispeed or pitch-preserving Lengthen. |
| V1_speed_ratio | 1.00 | Voice 1 transformation ratio. |
| V2_speed_ratio | 1.059 | Voice 2 transformation ratio. |
| Quantize_entries | Off | Choose BPM/note arithmetic instead of manual seconds. |
| Tempo_bpm | 120 | Reference tempo used for entry arithmetic and visualization beat grid. |
| Note_value | Quarter | Beat multiplier when quantized entry mode is active. |
| Entry_delay_s | 3.0 | Manual inter-voice entry delay. |
| Draw_visualization | On | Draw the 8×8 process visualization. |
| Play_output | On | Play final result. |
Custom Voice Details dialog
| Parameter | Default |
|---|---|
| V3_speed_ratio | 0.50 |
| V4_speed_ratio | 1.50 |
| V1_amplitude | 1.00 |
| V2_amplitude | 0.85 |
| V3_amplitude | 0.75 |
| V4_amplitude | 0.65 |
| V1_pan | −0.35 |
| V2_pan | +0.40 |
| V3_pan | −0.75 |
| V4_pan | +0.75 |
Speed ratios are clamped to 0.05–8.0; amplitudes to ≥0; pans to −1…+1.
Duration, normalization, silence trim & final fade
Measured voice body duration
All transformed chunks within one voice have the same duration. With a 40% overlap:
The pre-mix master duration is based on the latest actual voice end:
This v2.6 calculation prevents slower voices from being truncated merely because Voice 1 is shorter.
Peak normalization
After stereo mixing, the script measures the Sinc70 absolute peak. For any non-silent result it calls:
This is target peak normalization, not attenuation-only safety limiting. A quiet non-zero mix can be amplified to a 0.95 Sinc70 peak.
Automatic silence trim
The normalized stereo result is folded to mono for a 50 ms amplitude scan. A block counts as active if:
The script keeps approximately 50 ms of margin around the first and last detected active block. It performs the extraction only when more than 100 ms of silence would actually be removed from the beginning or end.
If the entire result remains below the threshold, trimming is skipped.
Final fade-out
After trimming, a raised-cosine fade to exact silence is applied over the last:
There is no second peak-normalization pass after trim or fade. The delivered output can therefore peak below 0.95 if the trimmed-away region contained the previous peak or if the final fade attenuates it.
Visualization
The current v2.6 Picture view follows the actual recomposition process:
- Chunk shuffle map: top source row shows chunks 1…N; each voice row shows the stored permutation that was actually rendered.
- Active span per voice: uses each voice's measured transformed OLA body duration, corrected for any leading trim offset.
- Left waveform: final channel 1.
- Right waveform: final channel 2.
- Summary strip: preset, source, voices, chunks, transform, fixed 40% overlap, voice ratios/pans, entry timing, output name and final duration.
Waveform scaling
Left and Right share the same amplitude range, calculated from the larger of the two channel peaks. Their displayed levels can therefore be compared directly.
Beat grid
Tempo_bpm; it does not mean that the actual entry delay was quantized.
Voice-entry markers
The colored dotted entry lines use the actual delayed entries minus any leading silence removed by auto-trim, so the markers remain aligned to the delivered output.
Output behavior
- Name:
<source>_poly_improv_v2. The historical suffix remains unchanged in v2.6. - Channels: always stereo.
- Source spatial image: not preserved; source is folded to mono before recomposition.
- Sample rate: source sample rate.
- Source object: unchanged.
- Minimum input duration: 1 second.
- Playback: optional through Play_output.