IRCAM RAVE Model — Neural Audio & Latent Transformations
Runs TorchScript RAVE/nn~ models offline from Praat, with model diagnostics, true latent-space transformations, optional controlled variation, sample-rate handling, and process visualisation.
What this does
IRCAM RAVE Model connects a selected Praat Sound to a TorchScript .ts model. The ordinary processing path calls the model's forward(). For compatible RAVE exports, a separate latent path calls encode(), applies a chosen transformation and optional energy layer, then calls decode().
encode() trajectory when the export exposes it. If a model has no usable encode(), the script falls back to a clearly labelled input feature proxy; the proxy is not presented as the model's latent space.Key features:
- Four actions — Process sound, Latent transform, Diagnose this model, Diagnose all models in folder.
- Eight latent operations + None — scale, offset, dimension gain, mute, isolate, smooth, freeze, jitter, or leave the encoded latent unchanged.
- Controlled latent energy — deterministic off, reproducible seeded variation, or a new variation on every run.
- Jitter or Smooth drift — absolute latent units or scaled relative to each encoded dimension's spread.
- Model-rate handling — resample to the declared model rate when available, optionally return to the input rate.
- Float audio interchange — Python writes 32-bit IEEE float WAV; normalisation is performed once.
- Model introspection — methods, nn~
*_paramsbuffers, declared sample rate, latent size and compression ratio.
Quick start
- For Process sound or Latent transform, select exactly one Sound in Praat.
- Run
IRCAM_rave_model.praat. - Choose a preset
.tsmodel or Custom, and confirm the models directory. - Choose the action. Process sound preserves the established forward-processing path. Latent transform opens a second dialog for latent controls.
- Leave Normalize = peak for the historical v1.2-style level, or choose none / rms when required.
- For deterministic latent processing leave Variation mode = Off. Use Seeded or New variation each run to inject controlled latent energy.
torch and numpy. The script probes available Python executables and reports a clear error if PyTorch cannot be found.Actions
| Action | Signal path | Use |
|---|---|---|
| Process sound | Sound → forward() → output | The established direct RAVE/nn~ processing path. This path remains unchanged by the latent-transform additions. |
| Latent transform | Sound → encode() → operation → optional energy → decode() → output | Direct manipulation of coordinates returned by the exported model's encode(). |
| Diagnose this model | No audio | Reports methods, parameter buffers, declared sample rate, latent size, compression and other exposed metadata. |
| Diagnose all models in folder | No audio | Runs the same capability probe for every .ts file in the chosen folder. |
Preset models & detected capabilities
The Praat menu provides seven preset filenames plus a Custom option. The table below summarises the RAVE exports identified by the diagnostics used to develop v1.3–v1.5. Custom models are probed at run time rather than assumed to match these capabilities.
| Preset | Latent | Compression | Declared rate | Notable capability |
|---|---|---|---|---|
break.ts | 8D | 2048 samples/frame | 44.1 kHz | encode, decode, prior |
darbouka_onnx.ts | 4D | 128 samples/frame | Not declared | High temporal latent resolution |
engine.ts | 4D | 2048 samples/frame | 44.1 kHz | encode, decode, prior |
InstantAlbania.ts | 8D | 2048 samples/frame | 44.1 kHz | encode, decode, latent-processing helpers |
isis.ts | 8D | 2048 samples/frame | 44.1 kHz | encode, decode, latent-processing helpers |
percussion.ts | 4D | 2048 samples/frame | 44.1 kHz | Mono input, native stereo output |
wheel.ts | 8D | 2048 samples/frame | 48 kHz | encode, decode, prior |
| Custom | Any user-specified .ts; capabilities are discovered rather than assumed. | |||
wheel.ts is a concrete reason to keep the default Resample to the model rate. For exports that do not declare a rate, the input is used as-is because the script does not guess.Latent transforms
The latent path operates on the coordinates exactly as returned by the exported model's encode(). Their statistical or semantic interpretation depends on that export.
| Operation | Rule | Main controls |
|---|---|---|
| Scale all dimensions | z' = amount × z | Amount |
| Offset | z'[d] = z[d] + amount | Dimension; 0 = all |
| Dimension gain | z'[d] = amount × z[d] | Dimension, Amount |
| Mute dimension | Sets one dimension to zero in the exported latent coordinate system. | Dimension |
| Isolate dimension | Keeps one dimension; all others are set to zero. | Dimension |
| Smooth over time | Moving average along time for every dimension, with edge padding. | Smoothing (ms) |
| Freeze one frame | Holds one encoded frame for a new requested duration. | Position 0–1, duration |
| Jitter | z' = z + amount × N(0,1) | Amount, Seed |
| None | Leaves the encoded latent unchanged. | Use this when you want only the Energy layer. |
Latent Energy — controlled instability
The Energy layer is applied after the deterministic latent operation and before decode():
| Control | Options | Meaning |
|---|---|---|
| Variation mode | Off / Seeded / New each run | Off is deterministic. Seeded is reproducible. New each run draws a new seed and reports it so the exact variation can later be reproduced in Seeded mode. |
| Energy type | Jitter / Smooth drift | Jitter is frame-to-frame instability. Drift low-pass-filters the noise for correlated motion. |
| Energy amount | 0–1 | Strength of injected variation. |
| Smoothness | 0–1 | Drift only. Maps to an approximate correlation time from 0.05 s to 5 s. |
| Energy target | All / One / High-variance | Chooses which encoded dimensions receive the added energy. |
| Energy scale | Relative / Absolute | Relative scales noise by each dimension's spread in the encoded input; Absolute uses latent-coordinate units directly. |
Reproducibility in v1.5.1
The Jitter operation and the Energy layer now use independent random streams. One user Seed still reproduces the complete result, but the two noise sources are no longer identical. A Fresh run reports its seed; entering that seed in Seeded mode reproduces the same Energy variation.
Parameters & advanced settings
Main dialog
| Parameter | Default | Description |
|---|---|---|
| RAVE model | break.ts | Seven preset filenames or Custom. |
| Action | Process sound | Forward processing, latent transform, or diagnostics. |
| Gain (dB) | 0 | Applied after normalisation. |
| Normalize | peak | none = model output untouched; peak = peak 0.99; rms = target −20 dBFS with peak guard. |
| Draw visualisation | yes | Draws the processing figure after audio actions. |
| Play result | yes | Plays the newly imported Sound. |
Advanced settings
| Setting | Options |
|---|---|
| Input shape | Auto-detect / BCT / B1T mono / CT |
| Output channels | Same as model / Force mono / Force stereo |
| SR handling | Resample to model rate (recommended) / Use as-is (v1.2 behaviour) |
| Output rate | Back to input rate / Keep model rate |
| Device | CPU / Auto (CUDA if available) / CUDA |
| Show latent | Enable/disable latent export for the figure |
| Output prefix | Default rave_out |
Visualisation
The v1.5.1 figure is designed to show what the algorithm actually did rather than a generic neural-audio illustration.
| Panel | Content |
|---|---|
| Title | Model, source, operation/energy information or normalisation summary, and detected model rate. |
| Spectrograms | Input and output on the same dB scale for meaningful comparison. |
| Waveforms | Input in grey and processed output in blue, on a shared amplitude/time display. |
| True latent trajectory | When encode() is available: up to the first eight dimensions through time, plus a z1×z2 path. In Latent transform mode, the encoded input is shown in grey and the transformed/energised latent in colour. |
| Feature proxy fallback | If no usable latent export exists, the panel is explicitly labelled Input feature proxy (NOT the model latent space). |
| Summary | Methods, latent size, device, sample-rate handling, level processing, channel/rate result and latent/energy description. |
Diagnostics
Diagnostics require no selected Sound. They load the model and report what the TorchScript export actually exposes, including:
- method names such as
forward,encode,decode,prior; - nn~ buffers such as
encode_params,decode_paramsandforward_params; - input/output channel counts and rate ratios from four-value method parameter buffers;
- declared sample rate when exposed as an attribute or buffer;
- latent dimension count and compression ratio when
encode_paramsis available.
Technical notes
Normalisation
- none — no scaling or clipping in the Python normaliser. Because the result is float WAV, peaks above full scale can be retained, with a warning.
- peak — scales the result to peak 0.99.
- rms — targets −20 dBFS RMS, attenuating further when required to keep the peak at or below 0.99.
- Gain is applied after the chosen normalisation, so it remains audible.
Sample rate
If the model declares a sample rate and it differs from the Sound, the default path resamples the input to the model rate with a rational windowed-sinc resampler. The result can then be returned to the original input rate or kept at the model rate. If the export does not declare a rate, the script uses the input rate as-is rather than guessing.
Output format
The Python worker writes a 32-bit IEEE float WAV. This avoids the unnecessary PCM16 quantisation used in the older runner and allows Normalize = none to remain genuinely unscaled.
Troubleshooting
- PyTorch not found: install
torchandnumpyin the Python environment Praat can launch. - Model file not found: verify the Models directory and selected/custom filename.
- Latent transform rejected: run Diagnose this model; latent processing requires usable
encode()anddecode(). - CUDA requested but unavailable: choose CPU or Auto.
- Unexpected pitch/time behaviour: inspect the reported sample-rate handling, especially for models with a declared rate different from the input.
- Need the same generative result again: use Seeded variation with the seed reported by the earlier Fresh run.