Stereo Mosaic — User Guide

Builds a stereo collage from two or more selected Sounds. Each source is converted to mono, partitioned into regions, optionally transformed, assigned to either the left or right stream, concatenated within that stream, then combined with optional M/S width and cross-channel bleed processing.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 0.5 (2026) License: MIT License Repo: GitHub
Contents:

What this does

Stereo Mosaic turns multiple source Sounds into two independently assembled mono streams that become the final left and right channels.

multiple selected Sounds → resample to highest source sample rate → mono work copies → region partitioning → optional duration scaling → L/R assignment → optional pitch / filtering / reverse / level shaping → concatenate assigned regions within L and R streams → pad shorter stream → optional M/S width → optional cross-channel bleed → combine to stereo → target peak 0.99
Channel assignment chooses which output stream receives each region. The script does not place every region on one common global timeline with complementary silence in the opposite channel.

Quick start

  1. Select at least two Sound objects.
  2. Run Stereo_Mosaic.praat.
  3. Choose Custom or one of the six named presets.
  4. Set Regions_per_file and the channel-assignment strategy.
  5. For Custom, adjust source-region overlap, output-stream gaps, pitch range, reverse probability and stereo width.
  6. Run. The result name includes the number of files, regions, preset and strategy.
Named presets also configure several internal transformation controls that are not exposed in the compact public form.

Source preparation

Common sample rate

The script finds the highest sample rate among all selected Sounds. Every lower-rate work copy is resampled to that rate with Praat quality 50.

Mono source bank

Every source is processed as mono:

Private work copies are shifted to a zero-based time domain before 0…duration region extraction.

The original stereo or multichannel spatial image of each source is not preserved. The final stereo field is constructed anew from region assignment and later width/bleed processing.

Region partitioning

Each file is processed independently. A preset may first choose one random positive source-time offset; Custom uses zero offset.

Overlap geometry

overlapFactor = 1 - Overlap_percent / 100 partitionDenominator = 1 + (Regions_per_file - 1) × overlapFactor regionDuration = (sourceDuration - timeOffset) / partitionDenominator regionStart(k) = timeOffset + (k - 1) × regionDuration × overlapFactor

Overlap_percent is internally clamped to 0…50%. Without duration variation, this geometry makes the final region end exactly at the end of the available source interval.

Region overlap refers to overlap between source extraction windows. It does not cause those regions to overlap in the output stream; output assembly is concatenative.

Preset duration variation

Some named presets independently vary each region's duration around the base region duration. The varied end is clipped to the source boundary. The start grid itself is still derived from the unvaried partition.

Random source offset

Presets that enable this feature draw one offset per file:

timeOffset ~ Uniform( 0, sourceDuration × random_time_offset_percent / 100 )

The complete partition is then fitted into the source interval from that offset to the file end.

Channel-assignment strategies

Strategies operate on the region number within each source file, except Random and Spiral which also introduce their own logic.

StrategyAssignment rule
Alternating regionsOdd region numbers → L; even region numbers → R. The pattern restarts for every source file.
Left first / Right secondRegion is L when region ≤ Regions_per_file / 2; all later regions are R. With an odd region count, the two groups are not necessarily equal.
Random splitIndependent 50/50 L/R draw for every region.
Reverse orderOdd region numbers → R; even region numbers → L. This reverses the Alternating channel assignment; it does not reverse temporal region order.
Inside outAssignment alternates according to the integer distance from the file's region midpoint.
Spiral patternDeterministic assignment from (fileIndex × 1.618 + regionIndex) mod 2.

Pan jitter used by presets

Some presets apply an additional stochastic L↔R reassignment after the main strategy. The internal control does not perform continuous panning.

p = pan_jitter_percent / 100 panShift ~ Uniform(-p, +p)

An L region can flip to R only when panShift > 0, followed by a second random test against panShift. An R region uses the symmetric negative case.

The numeric pan-jitter value is therefore not a literal region-flip percentage. It controls the range of the intermediate random variable.

Region transformations

After extraction, each valid region can pass through the following stages.

1. Duration scaling

Presets can draw a duration factor and use:

Lengthen (overlap-add): 75, 600, tempoScale

This is Praat overlap-add time scaling. Custom leaves this stage at 100%.

2. Pitch shift

Pitch is shifted with Praat Manipulation and a modified PitchTier:

pitchFactor = 2^(semitones / 12) PitchTier = originalPitchTier × pitchFactor

The segment is then resynthesized with overlap-add. This is pitch-tier resynthesis rather than varispeed source reading.

In Custom, Pitch_shift_semitones is treated as a magnitude: both +X and −X create the same random range −|X|…+|X|.

3. Channel-dependent filters

Only named presets can activate these hidden controls. The current Spectral Dance preset uses:

The filters use Praat Filter (pass Hann band) with a 100 Hz smoothing width.

4. Reverse

When Reverse_percent is active, each region receives an independent 0…100 random draw and is reversed when the draw falls below the requested threshold.

5. Preset amplitude variation

Amplitude variation is implemented as a random target peak, not multiplication of the existing level:

ampMult ~ Uniform( amplitude_variation_min, amplitude_variation_max ) / 100 Scale peak: ampMult × 0.95

After that, every region is divided by the preset's attenuation divisor.

6. Edge fades

Every processed region receives linear edge fades:

effectiveFade = min( fade_time_s, processedSegmentDuration / 2 )

The attack rises linearly from zero and the release falls linearly to zero. This is edge tapering, not a crossfade between adjacent output regions.

Independent L/R stream assembly

Regions are visited file by file and region by region. After processing, each region is appended only to the accumulator chosen by its final channel assignment:

if assigned L: Lstream = concatenate(Lstream, region) if assigned R: Rstream = concatenate(Rstream, region)

There is no placeholder silence inserted into the opposite stream for that event.

Consequently, the nth L region and nth R region do not necessarily correspond to the same source event or the same original position in the processing sequence. Once the two streams are combined, independently accumulated regions can overlap in output time.

Gaps

If Gap_ms > 0, that amount of silence is appended to the same channel immediately after every assigned region, including the final assigned region in that stream.

Channel length

After all regions are assembled, the shorter stream is padded with silence to match the longer one. If a strategy produces no regions on one side, that side becomes a full-length silence stream.

outputDuration = max( assembledLeftDuration, assembledRightDuration )

The exact duration depends on region assignment, source durations, preset duration scaling, clipping of varied regions, and any gaps.

Stereo width & cross-channel bleed

M/S width

After L/R lengths are equalized, Stereo_width_percent is applied through mid/side algebra:

M = (L + R) / 2 S = (L - R) / 2 w = Stereo_width_percent / 100 L' = M + wS R' = M - wS
WidthMeaning
0%Both channels become the same mid signal.
100%The assembled L/R streams are unchanged.
>100%The side component is amplified.

The public field is not internally clamped, so Custom values outside the usual 0…200% range extrapolate the same formula.

Cross-channel bleed

Some presets add a scaled copy of the opposite channel:

L'' = L' + bleed × R' R'' = R' + bleed × L'

The two additions use preserved pre-bleed copies, so the second assignment does not recursively feed the already-modified first channel.

Presets

Named presets overwrite both the visible controls and their internal transformation settings.

PresetRegions / strategyVisible controlsImportant internal processing
Classic Ping Pong4 / AlternatingOverlap 0%, gap 0, pitch 0, reverse 0%, width 100%No time/pitch variation; fade 50 ms; attenuation ÷1.1.
Glitchy Scatter8 / RandomOverlap 0%, gap 50 ms, pitch field 6, reverse 40%, width 150%20% source offset; ±30% duration variation; OLA factor 70–150%; pitch −5…+7 st; pan jitter 30; bleed 10%; random peak target 60–140% ×0.95; ÷1.3.
Spectral Dance6 / SpiralOverlap 15%, gap 0, pitch field 5, reverse 0%, width 130%10% source offset; OLA 90–110%; pitch −7…+5 st; L HPF 300 Hz; R LPF 4000 Hz; pan jitter 20; bleed 5%; random peak target 80–120% ×0.95; ÷1.1.
Wide Stereo Field5 / AlternatingOverlap 10%, gap 0, pitch 0, reverse 0%, width 180%Pan jitter 50; fade 60 ms; no bleed; no duration/pitch variation.
Dense Overlap12 / RandomOverlap 40%, gap 0, pitch field 3, reverse 25%, width 120%15% source offset; ±20% duration variation; OLA 85–115%; pitch −3…+3 st; pan jitter 40; bleed 15%; random peak target 70–130% ×0.95; ÷1.4.
Minimal Sparse3 / SplitOverlap 0%, gap 300 ms, pitch 0, reverse 10%, width 100%5% source offset; ±10% duration variation; OLA 95–105%; pan jitter 10; random peak target 90–110% ×0.95.
The visible Pitch_shift_semitones value shown after selecting a named preset is not necessarily the exact random pitch range used internally by that preset. The actual preset ranges are listed above.

Public parameters

ParameterDefaultExact behavior
PresetCustomCustom plus six named configurations.
Regions_per_file4Positive integer; number of source regions requested for every selected file.
Channel_strategyAlternatingOne of the six L/R assignment rules above.
Overlap_percent0Source-extraction overlap; internally clamped to 0…50%.
Gap_ms0Positive values append silence after every assigned output-stream region; zero or negative values add no gap.
Pitch_shift_semitones0In Custom, magnitude of a symmetric random pitch range −|X|…+|X|.
Reverse_percent0Per-region random reversal threshold.
Stereo_width_percent100M/S side multiplier after assembly; no internal clamp.
Draw_visualizationOnDraw current region distribution, output waveform, file legend and summary.
Play_resultOnPlay the completed stereo Sound.

Visualization

The v0.5 Picture view contains:

  1. Left channel region map — one rectangle for every region ultimately assigned to L.
  2. Right channel region map — the same for R.
  3. Output stereo waveform — the completed normalized result.
  4. File color legend — one color per source file, with truncated real object names when necessary.
  5. Summary bar — preset, file/region counts, strategy, overlap, gap, reverse, width, bleed, output duration, RMS and sample rate.

Reading the region maps

The x-axis is sequential position inside that output channel's region stream. The y-axis is the original source-file number. Rectangle color identifies the source file.

The maps do not use real output-time widths. Every region is drawn as one equal-width cell regardless of its processed duration.

The current v0.5 map also does not display the stored pitch-shift or reversal values as markers. Those transformations affect the audio but are not encoded in the region rectangles.

Output waveform

The stereo result is drawn with Praat's automatic waveform range. The panel is a view of the final result, not a calibrated comparison against the source files.

Output behavior

Final peak scaling is performed once on the combined stereo Sound so the L/R balance created by width and bleed processing is retained.