Wave Gesture Path Performer — Wearable Timbral-Path Performance
Turn a folder of sounds into a gesturally playable timbral path. The script orders the corpus by MFCC similarity, captures Tilt, Pan, and Roll from a Genki Wave ring over BLE-MIDI, and renders the recorded gesture as an overlapping polyphonic traversal through the corpus.
What this does
Wave Gesture Path Performer transforms a folder of audio files into a performable timbral space. Rather than assigning gestures directly to arbitrary file numbers, the script first analyses the corpus with MFCCs and constructs a nearest-neighbour path through timbrally related sounds. A recorded Wave-ring gesture then moves through that ordered path and controls how successive voices are selected and shaped.
Key Features:
- Wearable gesture control — Genki Wave ring via BLE-MIDI.
- MFCC timbral ordering — the corpus is reorganised as a nearest-neighbour path before performance.
- Three independent gesture streams — Tilt, Pan, and Roll.
- Global + local navigation — Tilt selects a location on the path; Pan scrubs locally around that location.
- Dynamic amplitude control — Roll controls the level of each triggered voice.
- Polyphonic overlap-add rendering — triggered sounds can overlap instead of cutting each other off.
- Auto-fit or absolute Tilt mapping — use the full corpus or retain a fixed physical 0-1 mapping.
- Multi-format corpus loading — WAV, AIFF/AIF/AIFC, FLAC, and MP3.
- MIDI activity diagnostics — the Python helper reports the CCs, ranges, counts, and channels actually received.
- Explanatory Praat Picture visualisation — raw gesture, mapped corpus path, and rendered waveform in one display.
- Offline rendering — the gesture is captured in real time, but the final audio is rendered reproducibly in Praat after the take.
Quick start
- Pair the Genki Wave ring as a BLE-MIDI device and make sure it appears as a MIDI input port.
- Install the Python MIDI dependencies:
midoandpython-rtmidi. - Place
wave_capture.pyinplugin_AudioTools/py/. - Run
Wave_Gesture_Path_Performer.praat. - Choose a folder containing at least two supported audio files.
- For a first run, keep Tilt mapping = Auto-fit take to full corpus, Grain seconds = 0.25, Voice seconds = 1.5, and Pan scrub amount = 4.
- Click Start. After the countdown, move the ring until the low end-beep.
- The rendered take appears in the Praat Objects list as
Wave_Take_1,Wave_Take_2, and so on.
How it works
1. The corpus is loaded and normalised
The script scans the selected folder for supported audio files, loads each file as a Praat Sound, converts multichannel files to mono, and resamples all successfully loaded sounds to the sample rate of the first valid corpus item.
2. MFCC mean vectors describe timbral similarity
Each sound is analysed with 12 MFCC coefficients using a 15 ms analysis window and 5 ms time step. The mean value of each coefficient across time becomes the sound's compact timbral descriptor.
3. A nearest-neighbour timbral path is constructed
Euclidean distances are calculated between the MFCC mean vectors. Starting from the first loaded sound, the script repeatedly chooses the nearest unvisited sound until every corpus item has been placed on a single ordered path.
4. Python captures the wearable gesture
Praat launches wave_capture.py, which opens the Wave MIDI port, records the configured CC streams at approximately 100 Hz, and writes a temporary tab-separated gesture file containing time, tilt, pan, and roll.
5. Praat renders the take
At each trigger time, the gesture is sampled, mapped to a corpus-path position, and used to select a source sound. A Hanning-windowed excerpt is extracted and summed into an output buffer. Overlapping excerpts form a polyphonic texture whose trajectory is determined by the recorded hand movement.
Gesture mapping
Tilt — global path position
Tilt chooses the main location along the MFCC-ordered corpus path.
Musical role: large-scale movement through timbral space.
Pan gesture — local scrub
Pan adds a positive or negative offset around the Tilt-selected position. The amount is limited by Pan scrub amount.
Musical role: local deviation, reversal, and non-linear movement around the current timbral region.
Roll — velocity / volume
Roll is mapped to voice amplitude between the internal minimum level and full scale.
Musical role: dynamic articulation and emphasis.
Controls
| Control | Default | Function |
|---|---|---|
| Folder | Blank | Corpus folder. Leave blank to choose a folder with a dialog. |
| Record seconds | 8.0 s | Duration of each captured gesture take. |
| Countdown seconds | 3 | Preparation time before recording begins. Audible cues mark countdown, start, and finish. |
| Grain seconds | 0.25 s | Time between triggers. Smaller values create denser activity and more frequent corpus changes. |
| Voice seconds | 1.5 s | Maximum duration of each triggered excerpt, limited by the duration of its source Sound. |
| Max voices | 8 | Maximum number of simultaneously active voices before oldest-voice stealing is applied. |
| Pan scrub amount | 4 positions | Maximum local path displacement created by the Pan gesture. |
| Tilt mapping | Auto-fit take to full corpus | Chooses between take-relative full-range mapping and fixed absolute 0-1 mapping. |
| Auto play | On | Plays the rendered take after it is created. |
Tilt mapping
Auto-fit take to full corpus
The script measures the take's own Tilt range using the 2nd and 98th percentiles and stretches that robust range across corpus positions 1...n.
Use: ensures that a natural hand gesture can access the full timbral path even when the physical sensor does not use the complete MIDI range.
Absolute (0-1)
The raw normalized Tilt value is mapped directly onto the corpus path.
Use: stable physical correspondence across takes, useful when the Wave already covers most of its MIDI range or when repeatable positions matter more than full-range access.
Corpus loading
Version 1.8 explicitly validates the corpus before any gesture is recorded.
| Behavior | Details |
|---|---|
| Supported formats | .wav, .aif, .aiff, .aifc, .flac, and .mp3, case-insensitive. |
| Folder depth | Only files directly inside the selected folder are searched. Subfolders are counted and reported but are not traversed. |
| Load validation | Every file is checked to confirm that a new Sound object was actually created. Failed loads are reported instead of accidentally reusing the previously selected Sound. |
| Minimum corpus size | At least 2 successfully loaded sounds are required. The script stops before recording if a performable path cannot be constructed. |
| Sample rate | All corpus Sounds are converted to the sample rate of the first successfully loaded item. |
| Channels | Multichannel corpus files are converted to mono for path analysis and rendering. |
Visualisation
After each take, the script produces a Praat Picture display organised as a compact explanation of the performance:
Gesture
Shows normalized Tilt, Pan, and Roll over time. In Auto-fit mode, the shaded region identifies the robust Tilt range fitted to the corpus.
Path
Shows the corpus position selected at every trigger. Grey indicates the position derived from Tilt alone; orange shows the position actually played after Pan scrub. Dot size reflects Roll-controlled volume.
Rendered take
Shows the final waveform together with trigger onsets, making the relationship between gesture density and resulting sound explicit.
MIDI diagnostics
The Python helper records all incoming MIDI activity during each take, not only the three configured control streams. The Info window therefore shows which CC numbers actually arrived, their value ranges, message counts, and MIDI channels.
Detected CC activity (this take): CC 2: 25-126 (357 messages, ch 1) <- Pan CC 1: 0-126 (317 messages, ch 1) <- Tilt CC 3: 0-126 (239 messages, ch 1) <- Roll
If a configured control shows too few messages or too little movement, the helper issues a warning and reports other active CCs. It does not remap controls automatically; the user remains in control of the Wave / Softwave configuration.
Technical behavior
- Uses 12 MFCC coefficients per corpus sound, averaged across analysis frames.
- Constructs a greedy nearest-neighbour path using Euclidean distance in MFCC mean-vector space.
- Captures Wave CC data at approximately 100 Hz through
mido/python-rtmidi. - Python-side MIDI port auto-detection prefers names containing Wave, Genki, Bluetooth, or BLE; an exact or partial port override can be set in the script.
- At each trigger, the nearest captured gesture sample to the trigger centre is used.
- Tilt determines the base corpus-path position; Pan adds a bounded integer scrub offset; Roll scales amplitude.
- Voice excerpts use Hanning windows and are summed into a mono output buffer.
- If the voice pool is full, the oldest active voice is stolen and a short taper is applied around the steal point.
- Source excerpt position moves through each chosen file according to performance time rather than always starting at the beginning.
- The rendered take is trimmed for trailing silence, peak-scaled to 0.99, and retained as a Praat Sound object.
- Temporary CSV, log, and done-sentinel files are cleaned automatically after use.
- The original corpus Sounds are removed from the Objects list after the session; rendered
Wave_Take_Nobjects remain.
Requirements & installation
mido and python-rtmidi.Install with:
python -m pip install mido python-rtmidi
| Component | Requirement |
|---|---|
| Praat | Praat with scripting support for external system calls. Praat 7 may ask for full trust because the tool launches Python and creates/deletes temporary files. |
| Python | Python 3. The frontend uses the library's OS-specific Python discovery convention. |
| Python helper | Place wave_capture.py in plugin_AudioTools/py/. The script also accepts the helper inside the selected corpus folder as a fallback. |
| Genki Wave | Wave ring paired as a BLE-MIDI input device. Default gesture CCs are Tilt = 1, Pan = 2, Roll = 3. |
| Softwave / MIDI configuration | The ring must send the intended movement streams as MIDI CC data. The helper's activity report can be used to verify the actual CC numbers and ranges. |
Limitations
- Top-level folders only: corpus files inside subfolders are not searched.
- Mono rendering: corpus Sounds are converted to mono and the generated take is mono.
- Path quality depends on the descriptor: the MFCC mean-vector path represents broad timbral similarity, not every perceptually relevant property of the source sounds.
- Greedy ordering: the nearest-neighbour path is locally constructed and is not a globally optimal travelling path through the corpus.
- Auto-fit is take-relative: the same physical Tilt can map to different corpus positions in different takes because each take's robust range is fitted independently.
- Pan is intentionally local: Pan scrub cannot replace global Tilt navigation; its range is limited by Pan scrub amount.
- Gesture data depends on MIDI configuration: a silent or incorrectly configured CC stream cannot be reconstructed by the Praat renderer.
- Dense triggering can mask corpus identity: short grain intervals and long overlapping voices can produce blended textures in which individual source changes become less obvious.
Outputs
Each completed take remains in the Praat Objects list as:
Wave_Take_1 Wave_Take_2 Wave_Take_3 ...
The script does not save an audio file automatically. The user can audition, inspect, rename, edit, or save any take from Praat after the session.
Applications
Embodied corpus navigation
Use case: perform a corpus by moving through a timbrally ordered path rather than selecting files manually or triggering fixed pads.
Gestural timbre improvisation
Use case: use broad Tilt motion for structural traversal while Pan creates local detours and Roll shapes articulation and dynamics.
Performance-derived offline composition
Use case: record physically expressive takes, then treat the resulting Wave_Take_N objects as fixed compositional material for further Praat AudioTools processing.
Timbral-path exploration
Use case: use the body as an exploratory interface for hearing how an MFCC similarity ordering behaves across heterogeneous instrumental, environmental, percussive, or textural corpora.
Research on embodied sound control
Use case: study the relationship between wearable gesture, corpus-space navigation, and rendered sound while retaining the captured MIDI trajectory and an explicit visual account of the mapping.
Workflow: heterogeneous corpus → gestural timbral traversal
Corpus: 10-30 contrasting short sounds.
Settings: Auto-fit Tilt, Grain 0.25 s, Voice 1.5 s, Pan scrub ±4 positions.
Performance: use Tilt for broad travel, Pan for local perturbation, and Roll for dynamic emphasis.
Result: an overlapping offline-rendered texture whose succession follows a recorded embodied path through MFCC-organised sound space.