Rhythmic Voice Flattener — User Guide
Compositional Voice Transformer. Flattens pitch to a centre frequency and rebuilds the audio as a formally shaped composition where musical interest arises from rhythm, timing, silence, density, grouping, accent, and gesture structure – not melody.
What this does
Rhythmic Voice Flattener takes a spoken or vocal recording and strips away its melodic contour, flattening every voiced segment to a single target pitch (40th percentile of the original F0). All musical interest is then constructed from rhythm, timing, silence, density, grouping, accent, and gesture structure. The result is a new composition built entirely from the source material, shaped by a formal design (arc, arrangement, rhythm mode) and a constraint‑satisfaction engine that assigns operations to each event.
- CSP rhythmic engine – backtracking solver enforces per‑gesture duration‑budget, density, and entropy constraints, replacing weighted‑table sampling from v2.
- Polytonal metric grid – each gesture tracks up to 3 simultaneous BPM frameworks; beat‑fitting searches all frameworks.
- Tempo dramaturgy arc – full accelerando / ritardando / metric‑modulation / wave trajectory across gestures.
- Note‑duration fetcher – assigns articulation labels (staccatissimo → sostenuto) and
note_dur_sto every event. - Fixed free mode – now routes through CSP with near‑zero grid weight instead of bypassing and leaving ratios at 1.0.
Quick start
- In Praat, select exactly one Sound object (spoken or vocal).
- Run script… →
RhythmicVoiceFlattener.praat. - Choose a Preset (none, minimal, pulse, scattered, mechanical, decay, breath, ritual, stutter_loop, ghost).
- For custom mode (preset = none), set:
- Arrangement – original / spectral / reversed / fragmented / density / palindrome
- Tension_arc – rising / arch / falling / wave / plateau / fractal / staircase / pulse
- Rhythm – free / soft / hard
- Sparsity – 0 (dense) … 1 (sparse)
- Optionally set Target_pitch_hz (0 = auto from source).
- Click OK. Praat segments events, extracts features, runs Python CSP engine, and reconstructs the result as
RVF_originalname_preset(or arrangement name).
numpy only. The CSP engine is pure Python – no external solvers.
PSOLA pitch flattening is done in Praat; very short events (<25 ms) are passed through unchanged to avoid artefacts.
The 10 presets
| Preset | Arrangement | Arc | Rhythm | Sparsity | Description |
|---|---|---|---|---|---|
| none | – | – | – | – | No override – use all explicit flags |
| minimal | original | falling | free | 0.8 | Sparse, slow, lots of silence |
| pulse | original | arch | hard | 0.2 | Tight grid, medium density |
| scattered | fragmented | wave | soft | 0.65 | Fragmented arrangement, high sparsity, wave arc |
| mechanical | density | arch | hard | 0.15 | Hard rhythm, dense, low silence, arch arc |
| decay | reversed | falling | soft | 0.5 | Falling arc, reversed arrangement, decreasing density |
| breath | original | arch | free | 0.75 | High sparsity, arch arc, slow silences |
| ritual | palindrome | wave | soft | 0.4 | Palindrome arrangement, motif scatter |
| stutter_loop | spectral | pulse | hard | 0.3 | Stutter‑heavy, burst ops, dense, pulse arc |
| ghost | original | falling | free | 0.9 | Very low amplitude offsets, sparse, drop‑heavy |
Each preset also sets internal parameters: _silence_scale, _op_bias, _tempo_shape, and _tempo_intensity.
Constraint‑Satisfaction Rhythmic Engine (v3 core)
- duration_budget_ratio [lo, hi] – total played duration / total original duration must land in this window.
- density_quota [lo, hi] – fraction of non‑skipped events.
- entropy_target [lo, hi] – normalised Shannon entropy of op‑type distribution (ensures variety or uniformity).
Constraints are derived from the tension arc, rhythm mode, and sparsity:
- High tension → tighter budget window, higher density quota, higher entropy.
- Low tension → wider budget, lower density, lower entropy.
- rhythm=free → wide budget window; grid weight ≈ 0.01 (not bypass).
- rhythm=hard → tight budget window.
- sparsity shifts density quota toward 0 (more drops).
Operations
| Op | Effect |
|---|---|
| grid | Nearest beat‑aligned duration |
| sustain | Longer (≈1.8× original) |
| compress | Shorter (≈0.55× original) |
| stutter | Very short + echo row |
| burst | Accent hit (+3.5 dB) |
| drop | Skip event (silence) |
| long_short | Pair: long then short |
| short_long | Pair: short then long |
The solver uses forward assignment with arc‑consistency lookahead. If it fails after 120 backtracks, constraints are widened by 25 % and a second attempt is made. Deviation from original targets is logged in stats (csp_budget_dev, etc.).
Articulation labels (note‑duration fetcher)
After duration ratios are finalised, each event receives an articulation label based on its played duration relative to one beat of the primary BPM framework:
| Fraction of beat | Label |
|---|---|
| < 10 % | staccatissimo |
| 10–25 % | staccato |
| 25–50 % | mezzo-staccato |
| 50–75 % | portato |
| 75–110 % | tenuto |
| > 110 % | sostenuto |
Dropped events (skip=True) get articulation = rest and note_dur_s = 0. The score CSV includes these columns for every event.
Parameters & defaults
Preset
Overrides arrangement, arc, rhythm, sparsity, and internal parameters. If preset ≠ none, manual settings are ignored.
Manual controls (when preset = none)
| Parameter | Options | Default |
|---|---|---|
| Arrangement | original / spectral / reversed / fragmented / density / palindrome | spectral |
| Tension_arc | rising / arch / falling / wave / plateau / fractal / staircase / pulse | arch |
| Rhythm | free / soft / hard | soft |
| Sparsity | 0 (dense) … 1 (sparse) | 0.5 |
| Target_pitch_hz | 0 = auto (40th percentile voiced F0) | 0 |
| Allow_repetition | boolean | 1 |
Score CSV columns (v3)
step– event index in outputsource_start_s,source_end_s– original segment timesvoiced– 1/0target_pitch_hz– flattened pitch (0 if unvoiced)duration_ratio– played / originalamplitude_db_offset– accent boost/cutsilence_after_s– rest after eventskip– 1 = dropped event (silence inserted)note_dur_s– actual played duration (seconds)articulation– staccatissimo … sostenuto / rest
Visualization (Praat picture)
When Draw_visualization = 1, the script draws:
- Original waveform (grey).
- Composed waveform (teal).
- Tension arc – graphical representation of the selected arc (rising, arch, fractal, etc.).
- Summary panel with:
- Source name, preset, effective arrangement/arc/rhythm, target F0, output duration.
- Events, gestures, tempo, sparsity.
- CSP deviation statistics (budget, density, entropy).
FAQ / troubleshooting
Install: pip install numpy. On Windows, the script uses python (Python launcher).
Check the score CSV – if many events have skip=1, the density quota may be very low.
Increase sparsity (toward 0) to keep more events. Also verify that the source has enough voiced segments – the engine needs something to work with.
PSOLA quality depends on segment length. Very short events (<25 ms) are left unchanged. The algorithm computes a safe pitch floor (max of 3/segDur and tgtF0*0.4) and skips PSOLA if the floor would exceed tgtF0*0.85. If many events are skipped, increase Target_pitch_hz manually to a higher value (still within the speaker's range).
The backtracking solver is fast for typical gestures (2–30 events). If deviations are high, the constraints may be impossible to satisfy simultaneously. This is often intentional – the engine relaxes and logs the deviation. Large deviations are not an error; they indicate that the design pushed the material to its limits.
Each gesture's bpm_frameworks is derived from its primary BPM by simple integer ratios (3:2, 4:3, 2:1, etc.). The nearest_beat_dur() function searches all active frameworks and returns the best‑fitting subdivision. This creates subtle polyrhythmic tension across gestures.
The stats file contains target_f0, n_events, n_gestures, bpm, preset, and three CSP deviation values. These are used in the visualization summary.