Rhythmic Voice Flattener — User Guide

Compositional Voice Transformer. Flattens pitch to a centre frequency and rebuilds the audio as a formally shaped composition where musical interest arises from rhythm, timing, silence, density, grouping, accent, and gesture structure – not melody.

Author: Shai Cohen Affiliation: Department of Music, Bar‑Ilan University, Israel Version: 3.0 (2026) License: MIT License Repo: GitHub
Contents:

What this does

Rhythmic Voice Flattener takes a spoken or vocal recording and strips away its melodic contour, flattening every voiced segment to a single target pitch (40th percentile of the original F0). All musical interest is then constructed from rhythm, timing, silence, density, grouping, accent, and gesture structure. The result is a new composition built entirely from the source material, shaped by a formal design (arc, arrangement, rhythm mode) and a constraint‑satisfaction engine that assigns operations to each event.

v3 core ideas:
  • CSP rhythmic engine – backtracking solver enforces per‑gesture duration‑budget, density, and entropy constraints, replacing weighted‑table sampling from v2.
  • Polytonal metric grid – each gesture tracks up to 3 simultaneous BPM frameworks; beat‑fitting searches all frameworks.
  • Tempo dramaturgy arc – full accelerando / ritardando / metric‑modulation / wave trajectory across gestures.
  • Note‑duration fetcher – assigns articulation labels (staccatissimo → sostenuto) and note_dur_s to every event.
  • Fixed free mode – now routes through CSP with near‑zero grid weight instead of bypassing and leaving ratios at 1.0.

Quick start

  1. In Praat, select exactly one Sound object (spoken or vocal).
  2. Run script… → RhythmicVoiceFlattener.praat.
  3. Choose a Preset (none, minimal, pulse, scattered, mechanical, decay, breath, ritual, stutter_loop, ghost).
  4. For custom mode (preset = none), set:
    • Arrangement – original / spectral / reversed / fragmented / density / palindrome
    • Tension_arc – rising / arch / falling / wave / plateau / fractal / staircase / pulse
    • Rhythm – free / soft / hard
    • Sparsity – 0 (dense) … 1 (sparse)
  5. Optionally set Target_pitch_hz (0 = auto from source).
  6. Click OK. Praat segments events, extracts features, runs Python CSP engine, and reconstructs the result as RVF_originalname_preset (or arrangement name).
Tip: Start with minimal (sparse, slow, lots of silence) to hear the basic architecture, then try pulse (tight grid, hard rhythm) for a mechanical feel. The stutter_loop preset uses heavy stutter/burst operations.
Important: Python dependencies: numpy only. The CSP engine is pure Python – no external solvers. PSOLA pitch flattening is done in Praat; very short events (<25 ms) are passed through unchanged to avoid artefacts.

The 10 presets

PresetArrangementArcRhythmSparsityDescription
none––––No override – use all explicit flags
minimaloriginalfallingfree0.8Sparse, slow, lots of silence
pulseoriginalarchhard0.2Tight grid, medium density
scatteredfragmentedwavesoft0.65Fragmented arrangement, high sparsity, wave arc
mechanicaldensityarchhard0.15Hard rhythm, dense, low silence, arch arc
decayreversedfallingsoft0.5Falling arc, reversed arrangement, decreasing density
breathoriginalarchfree0.75High sparsity, arch arc, slow silences
ritualpalindromewavesoft0.4Palindrome arrangement, motif scatter
stutter_loopspectralpulsehard0.3Stutter‑heavy, burst ops, dense, pulse arc
ghostoriginalfallingfree0.9Very low amplitude offsets, sparse, drop‑heavy

Each preset also sets internal parameters: _silence_scale, _op_bias, _tempo_shape, and _tempo_intensity.

Constraint‑Satisfaction Rhythmic Engine (v3 core)

Instead of independent weighted‑table sampling, the CSP engine solves for an assignment of operations across all events in a gesture that simultaneously satisfies three per‑gesture constraints:

Constraints are derived from the tension arc, rhythm mode, and sparsity:

Operations

OpEffect
gridNearest beat‑aligned duration
sustainLonger (≈1.8× original)
compressShorter (≈0.55× original)
stutterVery short + echo row
burstAccent hit (+3.5 dB)
dropSkip event (silence)
long_shortPair: long then short
short_longPair: short then long

The solver uses forward assignment with arc‑consistency lookahead. If it fails after 120 backtracks, constraints are widened by 25 % and a second attempt is made. Deviation from original targets is logged in stats (csp_budget_dev, etc.).

Articulation labels (note‑duration fetcher)

After duration ratios are finalised, each event receives an articulation label based on its played duration relative to one beat of the primary BPM framework:

Fraction of beatLabel
< 10 %staccatissimo
10–25 %staccato
25–50 %mezzo-staccato
50–75 %portato
75–110 %tenuto
> 110 %sostenuto

Dropped events (skip=True) get articulation = rest and note_dur_s = 0. The score CSV includes these columns for every event.

Parameters & defaults

Preset

Overrides arrangement, arc, rhythm, sparsity, and internal parameters. If preset ≠ none, manual settings are ignored.

Manual controls (when preset = none)

ParameterOptionsDefault
Arrangementoriginal / spectral / reversed / fragmented / density / palindromespectral
Tension_arcrising / arch / falling / wave / plateau / fractal / staircase / pulsearch
Rhythmfree / soft / hardsoft
Sparsity0 (dense) … 1 (sparse)0.5
Target_pitch_hz0 = auto (40th percentile voiced F0)0
Allow_repetitionboolean1

Score CSV columns (v3)

Visualization (Praat picture)

When Draw_visualization = 1, the script draws:

Tip: The tension arc preview helps you understand how the global design shapes the piece. The CSP deviations tell you how well the solver met the intended targets (0 = perfect).

FAQ / troubleshooting

“Python not found” or missing numpy

Install: pip install numpy. On Windows, the script uses python (Python launcher).

Output is silent or contains only silence

Check the score CSV – if many events have skip=1, the density quota may be very low. Increase sparsity (toward 0) to keep more events. Also verify that the source has enough voiced segments – the engine needs something to work with.

Pitch flattening sounds rough / artefacts

PSOLA quality depends on segment length. Very short events (<25 ms) are left unchanged. The algorithm computes a safe pitch floor (max of 3/segDur and tgtF0*0.4) and skips PSOLA if the floor would exceed tgtF0*0.85. If many events are skipped, increase Target_pitch_hz manually to a higher value (still within the speaker's range).

CSP solver takes a long time or returns large deviations

The backtracking solver is fast for typical gestures (2–30 events). If deviations are high, the constraints may be impossible to satisfy simultaneously. This is often intentional – the engine relaxes and logs the deviation. Large deviations are not an error; they indicate that the design pushed the material to its limits.

Polytonal metric grid

Each gesture's bpm_frameworks is derived from its primary BPM by simple integer ratios (3:2, 4:3, 2:1, etc.). The nearest_beat_dur() function searches all active frameworks and returns the best‑fitting subdivision. This creates subtle polyrhythmic tension across gestures.

Stats file

The stats file contains target_f0, n_events, n_gestures, bpm, preset, and three CSP deviation values. These are used in the visualization summary.