8-Channel Time Polyphony
Creates eight PSOLA-resynthesized versions of one source, each with its own constant duration factor. The voices can begin together and diverge in time, or enter at different times so that all eight converge on a common ending.
What this does
The script takes one selected Sound, converts it to a zero-based mono working source, and creates eight independent PSOLA resyntheses. Each voice receives one constant duration factor for its entire length. The factors therefore create a field of simultaneously related but differently stretched versions rather than a continuously changing time-warp.
- Eight voices: one duration factor per channel.
- Pitch-oriented time scaling: Praat Manipulation + DurationTier + overlap-add resynthesis.
- Two temporal organizations: common onset or staggered entries with a common ending.
- Per-voice edge fades: applied before entry padding and routing.
- Measured outcomes: requested and achieved duration factors are both reported.
- Five output layouts: octophonic, stereo-pair stems, odd/even quad stems, four-channel fold-down, or stereo odd/even mix.
Quick start
- Select exactly one Sound in Praat.
- Run
8-Channels_Time_Polyphony.praat. - Choose a Preset, or use Custom and enter eight duration factors.
- Choose Alignment: Common onset or Staggered entries.
- Set the PSOLA pitch range so that it matches the source material.
- Choose an Output_format.
- Leave Draw_visualization enabled to inspect the temporal field.
- Click OK.
Duration-factor model
This is a duration factor, not a speed factor. For example, a 10-second source with factor 1.30 aims for approximately 13 seconds; factor 0.70 aims for approximately 7 seconds.
s with D_i = D / s_i. Time Polyphony uses the opposite convention: D_i = D × r_i.
Requested versus achieved factor
PSOLA does not necessarily land on the requested duration with mathematical exactness. The script therefore measures every rendered voice and reports:
The visualization uses the requested factors to describe the compositional setup, while the drift and span displays use the measured voice durations.
Alignment: divergence and convergence
Common onset
All voices begin at time 0. Because their duration factors differ, corresponding source positions separate progressively in output time.
With different constant factors, common-onset voices can only diverge. Short voices finish earlier; long voices continue beyond them.
Staggered entries
The longest rendered voice defines the common endpoint. Every shorter voice receives silence before its onset:
The entries fan out, but all voices finish together. This produces genuine temporal convergence using constant duration factors.
Presets
The presets replace the eight duration factors. Diverging also forces Common onset; Converging forces Staggered entries. Other presets retain the Alignment selection from the form.
| Preset | Ch1–Ch8 duration factors | Structure |
|---|---|---|
| Custom | 1.00, 1.15, 0.85, 1.30, 0.70, 1.10, 0.90, 1.20 by default | User-controlled values. |
| Classic Polyphony | 1.00, 1.15, 0.85, 1.30, 0.70, 1.10, 0.90, 1.20 | Mixed expansion and contraction. |
| Slow Motion | 1.50, 1.70, 1.30, 1.60, 1.40, 1.80, 1.20, 1.90 | All voices longer than the source. |
| Irregular Fast Field | 0.50, 0.60, 0.40, 0.70, 0.30, 0.80, 0.25, 0.90 | All voices shorter than the source. |
| Dual-Rate Field | 1.00, 0.50, 1.00, 0.50, 1.00, 0.50, 1.00, 0.50 | Eight channels but only two distinct temporal variants. |
| Subtle Variation | 1.00, 1.05, 0.98, 1.02, 0.95, 1.03, 0.97, 1.01 | Close-duration field around unison. |
| Extreme Stretch | 3.00, 2.50, 3.50, 2.00, 4.00, 2.20, 3.80, 2.70 | Large PSOLA expansions. |
| Glitch Matrix | 0.15, 0.80, 0.30, 1.50, 0.20, 1.20, 0.40, 2.00 | Wide mixed factors; “Matrix” is only the preset title, not matrix processing. |
| Diverging | 0.70, 0.80, 0.90, 0.95, 1.05, 1.10, 1.20, 1.30 | Forces common onset; spread grows over time. |
| Converging | 0.70, 0.80, 0.90, 0.95, 1.05, 1.10, 1.20, 1.30 | Forces staggered entries; all voices finish together. |
| Unison | 1, 1, 1, 1, 1, 1, 1, 1 | One distinct temporal variant copied to eight channels. |
PSOLA resynthesis & edge fades
Processing chain for each voice
The two DurationTier points have the same value, so each voice uses one constant factor over the complete source.
Pitch-analysis range
Pitch_floor, Pitch_ceiling, and Analysis_time_step directly affect the Manipulation analysis and therefore PSOLA quality. The defaults—75 Hz, 600 Hz, 10 ms—are useful for many voices and speech sources but are not universal.
Per-voice edge fades
After resynthesis and before any entry padding, every voice receives optional raised-cosine fades using Praat's Fade in/out commands.
This placement is important. In common-onset mode, shorter voices may end in the middle of the final multichannel object, so a fade applied only to the completed output would not protect those internal endings. In staggered mode, all eight endings coincide, making edge control particularly important.
Parameters
| Parameter | Default | Actual role |
|---|---|---|
| Preset | Custom | Selects one of ten named factor sets or uses the eight Custom values. |
| Duration_factor_1 … 8 | 1.00, 1.15, 0.85, 1.30, 0.70, 1.10, 0.90, 1.20 | Constant duration multiplier per voice. Every value must be greater than zero. |
| Alignment | Common onset | Either all voices begin together, or shorter voices are delayed so all endings coincide. |
| End_fade | 0.010 s | Raised-cosine end fade per resynthesized voice; negative values become 0; capped to 10% of each voice. |
| Start_fade | 0.003 s | Raised-cosine onset fade per voice; negative values become 0; capped to 10% of each voice. |
| Pitch_floor | 75 Hz | Lower pitch bound for Manipulation analysis. |
| Pitch_ceiling | 600 Hz | Upper pitch bound; must be greater than Pitch_floor. |
| Analysis_time_step | 0.01 s | Time step used by To Manipulation. |
| Output_format | 8-channel octophonic | Selects the returned object/stem/downmix layout. |
| Scale_peak | 0.95 | Target for shared-gain normalization; invalid values are reset to 0.95. |
| Draw_visualization | on | Draws the v0.6 multi-panel process visualization. |
| Play_result | on | Plays the single output or a temporary odd/even stereo preview for multi-object stem formats. |
Output formats
Shared-gain stage
After the eight voices and any entry padding have been created, the largest absolute peak across all eight is measured. One common factor is applied to every voice:
This preserves the relative levels produced by the eight PSOLA renders. If every voice is effectively silent, normalization is skipped.
| Output format | Returned objects | Routing | Additional peak normalization |
|---|---|---|---|
| 8 channels — octophonic | 1 × 8-channel | out1–out8 = Ch1–Ch8 | No |
| 4 stereo pairs | 4 × stereo | Ch1|Ch2, Ch3|Ch4, Ch5|Ch6, Ch7|Ch8 | No |
| 2 quad groups | 2 × 4-channel | Odd = Ch1,3,5,7; Even = Ch2,4,6,8 | No |
| 4-channel fold-down | 1 × 4-channel | 1=Ch1+Ch2, 2=Ch3+Ch4, 3=Ch5+Ch6, 4=Ch7+Ch8 | Yes, final Scale_peak |
| Stereo mix | 1 × stereo | L = Ch1+Ch3+Ch5+Ch7; R = Ch2+Ch4+Ch6+Ch8 | Yes, final Scale_peak |
Why odd/even?
Several presets alternate factors between neighboring channel numbers. The odd/even grouping therefore separates temporal-rate families more clearly than a simple Ch1–4 / Ch5–8 split. In the Dual-Rate preset, for example, all odd channels are factor 1.0 and all even channels are factor 0.5.
Different voice lengths
Praat's channel combination pads shorter voices with silence, so a multichannel or stem object runs to the longest constituent voice. In staggered mode, all voices have already been padded at the beginning and finish together.
Preview playback
For the four-stereo-pair and two-quad-group formats, the script creates a temporary stereo monitor with odd voices on the left and even voices on the right, peak-normalizes it, plays it, and then removes it. The monitor is not one of the returned output objects.
Visualization
The v0.6 figure uses the suite-standard 8 × 8 layout and combines requested parameters with measured rendering outcomes.
Implementation notes & limits
- Exactly one Sound: the script checks the selection before showing the form.
- Mono-derived voices: multichannel input is converted to mono before PSOLA. Original stereo or surround imaging is not preserved.
- Non-zero start time: the working source is re-extracted when necessary so that its time domain begins at 0 and matches the DurationTier domain.
- Constant factor: the DurationTier contains identical values at the beginning and end; there is no time-varying duration curve.
- Preset names are descriptive: “Irregular Fast Field” is deterministic, and “Glitch Matrix” contains no matrix operation.
- PSOLA dependence: results depend on the suitability of pitch analysis for the source. This is not a general-purpose phase-vocoder or transient-preserving stretch engine.
- Factor validation: zero and negative duration factors are rejected.
- Longest output: for formats returning multiple objects, the report measures and prints the longest output-object duration rather than assuming one object represents the complete set.