Crossfade Concatenator — User Guide
Build a new mono or stereo Sound from whole files or extracted chunks, using explicit overlap-add crossfades, optional random ordering, and global or segment-based dynamics.
What this does
Advanced Concatenate with Crossfade v1.4 creates a new sequence from one or more selected Sound objects. Each source can contribute its whole duration or one or more fixed/random excerpts. The resulting chunks can remain in extraction order, be uniformly shuffled without replacement, or be sampled with replacement. Adjacent chunks are joined by a manual overlap-add crossfade using one of five selectable curves.
The processor supports mono and stereo material. Mixed sample rates are resampled to the sampling frequency of the first selected Sound. Mixed mono/stereo inputs are converted to the channel count of the first extracted chunk. Material with more than two channels is rejected.
Concatenate with overlap. The script applies one fade-out and one fade-in, then explicitly sums the two time-aligned buffers. This avoids a second hidden overlap window and makes the selected crossfade curve the only crossfade law applied.
Quick start
- Select at least one Sound object in Praat.
- Run
Concatenate_with_crossfade_v1.4.praat. - Choose a preset, or leave Custom and set chunking, order, crossfade, overlap, and dynamics manually.
- Enable Draw_visualization if you want the result waveform, segment map, dynamics envelope, file legend, and summary.
- Click OK. The final Sound is named
concat_crossfade_<PresetName>.
Parameters
Chunk extraction and order
| Parameter | Default | Implemented behavior |
|---|---|---|
| Preset | Custom | Selects one of six factory configurations or leaves the form values unchanged. |
| Chunk_mode | Whole file | Whole file, fixed-duration excerpt, or random-duration excerpt. |
| Fixed_chunk_duration_s | 2.0 s | Requested fixed excerpt length; capped at the source duration. |
| Min_chunk_duration_s | 0.5 s | Lower bound for random excerpt duration. |
| Max_chunk_duration_s | 3.0 s | Upper bound for random excerpt duration. If min > max, the script swaps them and reports the adjustment. |
| Chunks_per_file | 1 | Number of chunks extracted from every selected source. Whole-file mode therefore duplicates the whole file if this is greater than 1. |
| Randomize_order | On | Off = extraction order. On = random ordering. |
| Allow_repeats | Off | Only matters when randomization is on. Off uses a Fisher–Yates permutation; on samples each output position independently with replacement. |
Crossfade and overlap
| Parameter | Default | Implemented behavior |
|---|---|---|
| Crossfade_type | Linear | Linear, equal-power, S-curve, normalized exponential, or logarithmic. |
| Overlap_mode | Percentage | Percentage of incoming chunk, fixed duration, or random duration. |
| Overlap_percentage | 25% | Requested overlap as a percentage of the incoming chunk. |
| Fixed_overlap_s | 0.5 s | Requested fixed overlap. |
| Min_overlap_s | 0.1 s | Lower bound for random overlap. |
| Max_overlap_s | 1.0 s | Upper bound for random overlap. Inverted min/max values are automatically swapped. |
Dynamics and output
| Parameter | Default | Implemented behavior |
|---|---|---|
| Dynamics_mode | None | Flat, crescendo, diminuendo, swell, inverse swell, wave, random per segment, or terraced. |
| Wave_cycles | 2 | Number of sine-modulation cycles across the complete output in Wave mode. |
| Dynamics_depth_percent | 80% | Defines minAmp = 1 - depth. Values are clamped to 0–100%. |
| Scale_peak | 0.95 | For non-silent output, Praat Scale peak sets the final absolute peak to this value. Values above 1 are clamped to 1. |
| Draw_visualization | On | Draws the five-panel Picture report. |
| Play_result | On | Plays the final Sound after processing. |
Factory presets
| Preset | Chunking / order | Crossfade / overlap | Dynamics |
|---|---|---|---|
| Simple Crossfade | Whole file ×1; sequential; no repeats | Linear; 25% of incoming chunk | None |
| Smooth Collage | Random 1–4 s ×2/file; shuffled; no repeats | Equal-power; 30% | None |
| Rhythmic Chop | Fixed 0.5 s ×3/file; shuffled; no repeats | Linear; fixed 0.05 s | None |
| Cinematic Swell | Whole file ×1; sequential; no repeats | S-curve; 20% | Swell; 90% depth |
| Chaos Mix | Random 0.3–2.5 s ×3/file; random with repeats | S-curve; random 0.05–0.8 s | Random per segment; 60% depth |
| Granular Cloud | Random 0.05–0.3 s ×10/file; random with repeats | Equal-power; 50% | Wave; 3 cycles; 50% depth |
Chunk extraction, compatibility, and playback order
Extraction
Fixed and random chunks are taken from a uniformly random valid offset inside the source Sound. The code uses each Sound object's actual start and end times, so sources with non-zero time origins are handled correctly. A requested chunk longer than its source is shortened to the full available duration.
Sample rate and channels
The sampling frequency of the first selected Sound is the target. Any extracted chunk with a different rate is resampled to that rate with Praat's Resample: sr, 50.
Only mono and stereo chunks are accepted. The channel count of the first extracted chunk becomes the target channel format: stereo chunks are converted to mono when the target is mono; mono chunks are duplicated into stereo when the target is stereo.
Order
- Sequential: identity order 1, 2, …, N.
- Random, no repeats: an in-place Fisher–Yates shuffle; every extracted chunk appears exactly once.
- Random with repeats: every output slot independently draws an integer from 1…N. A chunk may appear several times and another may not appear at all.
Crossfade curves and overlap rules
For normalized overlap position u from 0 to 1, the outgoing and incoming chunks are multiplied by complementary curves, then summed by manual overlap-add.
| Type | Incoming gain | Outgoing gain | Property |
|---|---|---|---|
| Linear | u | 1-u | Amplitude weights sum to 1. |
| Equal-power | sqrt(u) | sqrt(1-u) | Squared gains sum to 1; intended to reduce the center power dip for uncorrelated material. |
| S-curve | 0.5 - 0.5 cos(pi u) | 0.5 + 0.5 cos(pi u) | Smooth complementary cosine curve. |
| Exponential | (1-exp(-4u))/(1-exp(-4)) | (exp(-4u)-exp(-4))/(1-exp(-4)) | Fast-start curve normalized exactly to 0/1 endpoints in v1.4. |
| Logarithmic | ln(1+9u)/ln(10) | 1 - incoming | Alternative fast-start complementary curve. |
Actual overlap duration
The selected overlap mode first produces a requested overlap. The script then caps it to 90% of both neighboring chunks:
capacity = min(0.9 × incomingDuration, 0.9 × outgoingDuration) overlap = min(requestedOverlap, capacity) minimum = min(0.01 s, capacity) overlap = max(overlap, minimum)
Therefore a join never consumes an entire neighboring chunk. When both chunks permit it, the overlap is at least 10 ms; for extremely short chunks, the available 90% capacity takes precedence.
Dynamics modes
Let d = Dynamics_depth_percent / 100 and minAmp = 1-d. A depth of 0% gives unity gain; 100% allows the chosen envelope to reach zero where its shape reaches its minimum.
| Mode | Implemented envelope |
|---|---|
| None | Unity gain. |
| Crescendo | Linear rise from minAmp at the beginning to 1 at the end. |
| Diminuendo | Linear fall from 1 to minAmp. |
| Swell | Triangular macro-envelope: minAmp → 1 → minAmp. |
| Inverse swell | 1 → minAmp → 1. |
| Wave | Sine modulation across the full result; starts at minAmp and completes Wave_cycles cycles. |
| Random per segment | Each playback position receives an independent gain uniformly drawn from [minAmp, 1]. |
| Terraced | Up to five equally spaced levels from minAmp to 1, repeated cyclically across segment positions. With only one total chunk, its gain is 1. |
Processing pipeline
- Store the selected Sound object IDs and use the first Sound's sampling frequency as the target.
- Apply the selected preset, then sanitize inverted chunk/overlap ranges, dynamics depth, and peak target.
- Extract
Chunks_per_filechunks from each source. - Resample mismatched chunks to the first Sound's sampling frequency.
- Reject channel counts other than 1 or 2, then convert all chunks to the first extracted chunk's mono/stereo format.
- Build sequential, shuffled, or with-replacement playback order.
- For Random/Terraced dynamics, assign the per-position gains before rendering.
- Starting from the first playback chunk, calculate each overlap, apply exactly one fade-out/fade-in pair, and sum the buffers with manual overlap-add.
- Apply any continuous whole-output dynamics envelope.
- If the result is non-silent, scale its absolute peak to
Scale_peak; if completely silent, skip peak scaling. - Rename to
concat_crossfade_<PresetName>, compute output RMS, optionally draw, clean temporary chunks/control objects, and optionally play.
Scale peak changes the level so the absolute peak equals the requested target. It can attenuate a hot render or amplify a quiet one.
Visualization
When Draw_visualization is enabled, v1.4 draws a vertically stacked 8-inch report:
- Header: preset, source count, chunk count, crossfade type, and dynamics mode.
- Panel A — Result waveform: final output with dashed segment-start markers.
- Panel B — Segment map: time-aligned colored rectangles; color identifies source file and the number shows playback position when space permits.
- Panel C — Dynamics envelope: flat unity line for None; analytic curve for continuous modes; for Random/Terraced, a parallel one-channel control signal is assembled through the same overlap-add crossfades as the audio.
- Panel D — File color legend: one swatch per selected source, with long names truncated for display.
- Panel E — Summary: preset/order, chunk mode, crossfade and overlap settings, dynamics, output name/duration, total overlap, output RMS, and sample rate.
sqrt(u) + sqrt(1-u) > 1 away from the endpoints; the power relation, not the amplitude sum, is constant.
Limits and edge cases
- Mono/stereo only: any extracted chunk with another channel count stops the script.
- Target format comes from the first source/chunk: sample rate follows the first selected Sound; mono/stereo format follows the first extracted chunk.
- Randomness is not seeded in the form: random chunk offsets, durations, orders, overlaps, and segment gains may change between runs.
- Overlap is constrained: every join is capped at 90% of each neighbor, regardless of a larger percentage/fixed/random request.
- Scale_peak above 1 is clamped: this avoids intentionally creating a post-normalization peak above digital full scale.
- Silent output: final peak normalization is skipped instead of calling
Scale peakon an all-zero Sound. - Random with repeats: the number of output positions remains
totalChunks, but the underlying extracted chunks are sampled with replacement.