Corpus Concatenative Codec — Neural Codec Corpus Synthesis

Corpus-based concatenative synthesis using a neural audio codec (EnCodec or DAC) as the matching token space. Four modes: Match Build corpus Draw Gesture rhyme. Draw now uses an interactive Praat Demo-window contour surface with corpus-brightness feedback, built-in gesture generators, audible preview, and reusable contours.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: Praat front-end 1.14.0 / Python backend 1.6 (2026) License: MIT License Repo: https://github.com/ShaiCohen-ops/Praat-plugin_AudioTools
Contents:

What this does

This script implements corpus-based concatenative synthesis using a neural audio codec (EnCodec or DAC) as the matching token space. It offers four modes: Match (synthesise from a selected Sound), Build corpus (encode a corpus once), Draw (compose a time-varying brightness path in an interactive Demo window and let the corpus voice it), and Gesture rhyme (re-voice an abstract kinetic gesture using codec-token transition structure).

What makes this different? The Gesture rhyme mode (new in v1.6) re-voices abstract kinetic gestures — accelerating clicks, a bouncing ball, an explosive attack decaying into hiss, a microtonal dive, tremolo flutter — using corpus grains whose codec-token TRANSITION structure (hashed bigrams) rhymes with the source, NOT whose timbre matches. It weights the bigram section high and the histogram/energy sections low, so a click-train can be voiced by speech syllables or field recordings that simply move the same way frame-to-frame. This is "compositional gestural rhyming" — the system searches for grains whose internal movement patterns match the source's kinetic shape.

Key Features:

Gesture rhyme vs. Match: Match mode finds grains whose timbre (token histogram) matches the source. Gesture rhyme finds grains whose movement (token bigram transitions) matches the source. A click-train and a speech syllable both have rapid frame-to-frame token changes; gesture rhyme can voice a click-train with speech syllables, field recordings, or instrument noises that simply move the same way. Literal sound identity is deliberately ignored.

Quick start

  1. Prepare a corpus audio folder (sounds to use as source material).
  2. First time: run Build / replace corpus — set Corpus_audio, choose DAC or EnCodec, and set the corpus grain/hop values. The analysed corpus is cached and reused.
  3. Choose a mode. Match and Gesture rhyme require a selected Sound object; Draw does not.
  4. For Draw, choose Interactive (Demo window) or Contour file. Interactive Draw opens the dedicated contour surface; a saved contour replays its stored time→brightness trajectory and duration.
  5. In Interactive Draw, click empty space to add points; click a point and then click elsewhere to move it; Shift-click a point to delete it. Use the on-screen generators/transforms, Preview (P), and Render (Enter).
  6. For Gesture rhyme, adjust Bigram_weight (high = kinetic matching), Hist_weight (low = de-emphasise timbre), and Sequence_context (context smoothing).
  7. The Python backend synthesises the result and Praat imports it as a new Sound. A rendered Draw run also saves the normalized contour actually used, so the gesture can be reproduced later.
Quick tip: Use Gesture rhyme with a high Bigram_weight (4.0–8.0) and low Hist_weight (0.1–0.5) to voice abstract gestures. Try a source that's just a click train or a bouncing ball — the corpus will voice it with grains that move the same way. Sequence_context (1–3) smooths the matching over preceding windows, privileging sustained transition structure over isolated windows.
Important: Python dependencies required: pip install numpy scipy soundfile. For EnCodec: pip install encodec. For DAC: pip install dac. The corpus index is built once and reused. Gesture rhyme requires an existing index — it never builds or reslices the corpus. The "mock" codec works without torch and is useful for testing.

4 Modes

Match (synthesise from corpus)

Select a Sound in Praat. The script detects onsets, subdivides into analysis windows, matches each sub-window against the corpus (timbre + loudness), and reconstructs the output.

Use: Transform any sound using a corpus of your choice.

Build corpus index

One-time encoding of a corpus folder. Slices audio into overlapping grains, encodes each into codec tokens, stores features and metadata.

Use: Prepare a corpus for Match, Draw, and Gesture rhyme modes.

Draw (brightness contour → corpus)

Draw and edit a time→brightness contour in a dedicated Praat Demo window. The display shows the corpus's own brightness distribution and previews which available brightness each grain-rate step will select.

Use: Compose a path through the corpus, audition it before rendering, and save the exact contour for later replay.

Gesture rhyme NEW

Re-voice an abstract kinetic gesture using codec-token transition (hashed-bigram) rhyming. High bigram weight, low histogram weight. A click-train can be voiced by speech syllables or field recordings that move the same way.

Use: Voice abstract kinetic shapes — accelerating clicks, bouncing balls, explosive attacks, tremolo flutter.

Gesture rhyme feature-compatibility: The stored corpus feature matrix is split into histogram and bigram sections. The script re-weights each section at match time (hist_weight, bigram_weight) — no corpus rebuild is needed. Sequence_context averages source features over preceding sub-windows, privileging sustained transition structure over isolated windows.

Draw Mode — Interactive Brightness-Contour Composition

Draw mode treats the analysed corpus as a navigable brightness field. Each corpus grain has a spectral-centroid value; the backend maps the corpus's log spectral-centroid range to a normalized 0..1 brightness axis (0 = darkest, 1 = brightest). The drawn contour is sampled at the chosen grain rate, and each step selects the corpus grain whose available brightness is nearest to the target, with repeat penalty applied during the actual render.

What the Draw window shows: the blue line is your gesture; the orange step line shows the nearest available grain brightness at each grain-rate step (a visual preview that does not include repeat penalty); background shading indicates corpus-grain density; red bands mark brightness regions for which the corpus contains no grains. A density strip at the right also shows the corpus's approximate minimum and maximum spectral centroid in Hz.

Editing the contour

ActionResult
Click empty spaceAdd a contour point.
Click a point, then click elsewhereMove the selected point. Interior points move in time and value; the two end points remain fixed in time and move vertically only.
Shift-click a pointDelete an interior point. End points cannot be deleted.
U / UndoUndo the last point edit, generator, or time transform.
R / ResetReturn to the default dark→bright ramp.
S / StretchToggle legacy min/max stretching of the drawn range to 0..1.
P / PreviewRender and play the current gesture using the same backend Draw call and synthesis parameters as the final render, then return to editing.
Enter / RenderRender the current gesture, import the result, and save the normalized contour actually used.
Esc / CancelLeave Draw without rendering.

Built-in gesture generators and time transforms

Sine, Triangle, and Angular rewrite the editable point list across the whole Draw duration. The cycles − / + controls set 1–32 cycles. The generator range is taken from the current drawing's lowest and highest values; if the drawing is essentially flat (range < 0.02), the generator uses the full 0..1 brightness range. Angular creates a new jittered zigzag variant each time it is invoked.

Same wave faster ×2 compresses the current drawing into half the duration and repeats it twice. Same wave slower ÷2 takes the first half of the current drawing and stretches it over the full duration. These operations modify the current contour and remain fully editable and undoable.

Absolute brightness vs. Stretch

By default, the vertical axis is absolute within the current corpus: a point at 0.25 targets the lower quarter of that corpus's normalized log-brightness range, while a point at 0.75 targets the upper quarter. Turning Stretch drawn range on reproduces the older RealTier behaviour: the minimum and maximum values in the current drawing are expanded to 0 and 1 before synthesis; a completely flat drawing maps to 0.5. When Stretch is active, the window also shows the effective stretched contour.

Preview, Render, and reusable contour files

Preview writes a temporary contour and calls the same Draw backend with the same grain rate, crossfade, repeat penalty, duration, and effective stretch result used by Render. It produces a temporary WAV, plays it in full, deletes it, and returns to the editor. Preview does not save a persistent contour copy.

Render writes the gesture in the compact human-readable corpus_draw_contour 1 format and asks the backend to save the normalized time→brightness trajectory actually used. The persistent copy is stored under Praat's preferences directory in corpus/draw_contours/. You can later choose Draw source = Contour file and load that file to reproduce the same gesture exactly; the contour file's stored duration is reused.

No RealTier editor in the current interactive workflow. Since Praat front-end v1.12.0, Interactive Draw is handled entirely in the Demo window. The Python backend still accepts the legacy --tier path for backward compatibility, but the current Praat front-end writes and uses explicit contour files.

4 Match Presets

PresetOnset Interval (ms)Analysis Grain (ms)Analysis Hop (ms)Energy WeightCrossfade (ms)Repeat PenaltyCharacter
Rhythmic4040201.5100.10Follows transients closely — fast, articulated.
Textural120120600.5600.02Smeared, washed — ambient texture.
Faithful6050251.0200.15Track source closely — balanced.
Sparse80100801.0150.30Distinct granular stutter — sparse.

Codec Token Space — Features for Matching

Token extraction

Each audio grain is encoded by a neural codec (EnCodec or DAC) into tokens: [n_codebooks, n_frames] of integer token IDs.

Token IDs are categorical symbols — we never take Euclidean distance on raw IDs.

Searchable features

  • Per-codebook token histograms — marginal distribution of each codebook (1024 bins) — timbre
  • Hashed bigram transitions — ordered token-pair structure kept separate per codebook; each codebook hashes transitions through a 64-bit avalanche mix into 1021 buckets — local sequence / kinetic structure
  • Features are L1-normalised and concatenated; compared with cosine distance.

Energy term

Cosine distance on token features captures timbre but not loudness. A near-silent grain and a loud one can have similar token-histogram shapes. The match function adds a separate log-RMS energy term: energy_dist = |log(RMS_corpus) - log(RMS_source)| / 4, weighted by energy_weight.

Gesture rhyme weights

  • Bigram_weight (default 4.0) — high = kinetic motion drives the match
  • Hist_weight (default 0.5) — low = literal timbre de-emphasised
  • Energy_weight (default 0.2) — low = energy can't dominate kinetic match
Why hashed bigrams? Bigram transitions capture how tokens change frame-to-frame — acceleration, deceleration, oscillation, abrupt shifts. A click-train and a speech syllable both have rapid frame-to-frame token changes; gesture rhyme finds grains whose internal movement patterns match the source's kinetic shape, regardless of what they actually sound like.

Gesture Rhyme Mode — Hashed-Bigram Kinetic Matching

Gesture rhyme workflow:
  1. Select a Sound object — the abstract gesture source. This can be a click train, a bouncing ball, an explosive attack decaying into hiss, a microtonal dive, tremolo flutter — anything with a kinetic shape.
  2. Choose Mode = Gesture rhyme.
  3. Set Bigram_weight (high = kinetic matching), Hist_weight (low = de-emphasise timbre), Energy_weight (low = energy de-emphasised).
  4. Set Sequence_context (0–3) to smooth matching over preceding sub-windows.
  5. Click OK — the system searches the EXISTING corpus for grains whose token-transition structure rhymes with the source.
  6. Output is voiced by corpus grains that move the same way, regardless of their literal sound.
Feature weighting:
  • Bigram_weight 4.0 — strong kinetic matching
  • Hist_weight 0.5 — timbre de-emphasised
  • Energy_weight 0.2 — loudness de-emphasised
  • Sequence_context 1–3 — average over preceding windows for smoother transitions
Kinetic shapes you can voice:
  • Accelerating clicks — rapid attack → voiced by grains with increasing token-change rate
  • Bouncing ball — decaying impacts → voiced by grains with damping structure
  • Explosive attack → hiss — sharp onset, noisy decay → voiced by grains with similar envelope
  • Microtonal dive — descending pitch glide → voiced by grains with descending token trajectory
  • Tremolo flutter — rapid amplitude modulation → voiced by grains with oscillating tokens
Compositional gestural rhyming: The source gesture is an ABSTRACT kinetic shape, not a sound you want to hear. The corpus provides the sonic material. The system finds grains whose internal movement (token transitions) rhymes with the gesture. A click-train can be voiced by speech syllables, field recordings, or instrument noises that simply move the same way. This is a new way to compose: design a gesture, then find corpus material that voices it.

Applications

Gesture rhyme — Voice abstract kinetic shapes

Use case: Design a kinetic gesture (accelerating clicks, bouncing ball, explosive attack, microtonal dive) and voice it with corpus grains that move the same way.

Settings: Gesture rhyme mode, Bigram_weight=4.0, Hist_weight=0.5, Sequence_context=1–3.

Match — Rhythmic corpus synthesis

Use case: Replace drum hits or percussive sounds with corpus grains that match the rhythm.

Settings: Rhythmic preset (short analysis windows, high energy weight).

Draw mode — Brightness contour composition

Use case: Compose a piece by drawing brightness contours, using the corpus as a timbral palette.

Settings: Draw mode, choose Interactive (Demo window), set Draw duration and Grain rate, then shape the contour manually or with Sine / Triangle / Angular generators. Use Preview before Render when auditioning live alternatives.

Workflow: Gesture rhyme — Click train → speech syllables

Gesture source: A click train (rapid, decaying clicks).
Corpus: Folder of speech syllables (vowels, consonants).
Settings: Gesture rhyme mode, Bigram_weight=4.0, Hist_weight=0.3.
Result: The output is voiced by speech syllables whose token-transition structure matches the click train — the rhythm and kinetic shape of the clicks, but filled with vocalic content.

Workflow: Gesture rhyme — Bouncing ball → field recordings

Gesture source: A bouncing ball (decaying impacts, decreasing intervals).
Corpus: Field recordings (doors, footsteps, water drops).
Settings: Gesture rhyme mode, Sequence_context=2.
Result: The output voices the bouncing ball with field recordings that share the same kinetic structure — decaying impacts, decreasing time intervals — but using environmental sounds.

Workflow: Draw mode — Brightness glissando

Corpus: Synthesiser patches (bright to dark).
Settings: Draw mode, Interactive (Demo window), 8 s duration, 100 ms grain rate.
Contour: Start from the default 0→1 ramp or draw a new path; the window immediately shows corpus density and the nearest available brightness sequence.
Result: The output travels from darker to brighter regions of the corpus — a corpus-driven spectral glissando. Use Preview to audition the current path before Render.

Troubleshooting:
• No corpus index found: Build corpus mode first, or run Match once to auto-build. Gesture rhyme never builds the corpus.
• Gesture rhyme output sounds like timbre matching: Increase Bigram_weight (8.0–12.0) and decrease Hist_weight (0.1–0.3). The match should be driven by bigram transitions, not histograms.
• Gesture rhyme output is jerky / discontiguous: Increase Sequence_context (1–3) to smooth the match over preceding windows. This privileges sustained transition structure.
• Codec not installed: For EnCodec: pip install encodec. For DAC: pip install dac. Use "mock" to test without torch.
• Draw window has red horizontal regions: those brightness ranges contain no corpus grains; redraw through denser regions or rebuild the corpus with more varied material.
• Draw reports that brightness data is missing: the corpus index was built by an older version; rebuild / replace the corpus so per-grain spectral-centroid data is stored.
• Preview or Render pauses Praat: this is expected while the backend renders (and while Preview plays the result). Preview uses the same synthesis path as Render. On Praat 7.0, file-writing / subprocess operations may also trigger the normal trust prompt.
• Closing the Draw Demo window: closes the interactive Draw session and stops the script; use Cancel if you want an explicit exit.

Visualisation & TextGrid

When Import_textgrid is enabled, all modes export a TextGrid with an IntervalTier named "grains". In Gesture rhyme mode, the metadata JSON includes gesture_contribution (bigram match), timbre_contribution (histogram match), and energy_contribution — so you can see exactly which component drove each selection.