IRCAM RAVE Model — Neural Audio & Latent Transformations

Runs TorchScript RAVE/nn~ models offline from Praat, with model diagnostics, true latent-space transformations, optional controlled variation, sample-rate handling, and process visualisation.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 1.5.1 (2026) License: MIT License Repo: Praat AudioTools
Contents

What this does

IRCAM RAVE Model connects a selected Praat Sound to a TorchScript .ts model. The ordinary processing path calls the model's forward(). For compatible RAVE exports, a separate latent path calls encode(), applies a chosen transformation and optional energy layer, then calls decode().

True latent data when available. Since v1.3 the figure uses the model's actual encode() trajectory when the export exposes it. If a model has no usable encode(), the script falls back to a clearly labelled input feature proxy; the proxy is not presented as the model's latent space.

Key features:

Quick start

  1. For Process sound or Latent transform, select exactly one Sound in Praat.
  2. Run IRCAM_rave_model.praat.
  3. Choose a preset .ts model or Custom, and confirm the models directory.
  4. Choose the action. Process sound preserves the established forward-processing path. Latent transform opens a second dialog for latent controls.
  5. Leave Normalize = peak for the historical v1.2-style level, or choose none / rms when required.
  6. For deterministic latent processing leave Variation mode = Off. Use Seeded or New variation each run to inject controlled latent energy.
Dependencies: Praat 6.3+ and Python 3 with torch and numpy. The script probes available Python executables and reports a clear error if PyTorch cannot be found.

Actions

ActionSignal pathUse
Process soundSound → forward() → outputThe established direct RAVE/nn~ processing path. This path remains unchanged by the latent-transform additions.
Latent transformSound → encode() → operation → optional energy → decode() → outputDirect manipulation of coordinates returned by the exported model's encode().
Diagnose this modelNo audioReports methods, parameter buffers, declared sample rate, latent size, compression and other exposed metadata.
Diagnose all models in folderNo audioRuns the same capability probe for every .ts file in the chosen folder.

Preset models & detected capabilities

The Praat menu provides seven preset filenames plus a Custom option. The table below summarises the RAVE exports identified by the diagnostics used to develop v1.3–v1.5. Custom models are probed at run time rather than assumed to match these capabilities.

PresetLatentCompressionDeclared rateNotable capability
break.ts8D2048 samples/frame44.1 kHzencode, decode, prior
darbouka_onnx.ts4D128 samples/frameNot declaredHigh temporal latent resolution
engine.ts4D2048 samples/frame44.1 kHzencode, decode, prior
InstantAlbania.ts8D2048 samples/frame44.1 kHzencode, decode, latent-processing helpers
isis.ts8D2048 samples/frame44.1 kHzencode, decode, latent-processing helpers
percussion.ts4D2048 samples/frame44.1 kHzMono input, native stereo output
wheel.ts8D2048 samples/frame48 kHzencode, decode, prior
CustomAny user-specified .ts; capabilities are discovered rather than assumed.
Sample-rate note: wheel.ts is a concrete reason to keep the default Resample to the model rate. For exports that do not declare a rate, the input is used as-is because the script does not guess.

Latent transforms

The latent path operates on the coordinates exactly as returned by the exported model's encode(). Their statistical or semantic interpretation depends on that export.

OperationRuleMain controls
Scale all dimensionsz' = amount × zAmount
Offsetz'[d] = z[d] + amountDimension; 0 = all
Dimension gainz'[d] = amount × z[d]Dimension, Amount
Mute dimensionSets one dimension to zero in the exported latent coordinate system.Dimension
Isolate dimensionKeeps one dimension; all others are set to zero.Dimension
Smooth over timeMoving average along time for every dimension, with edge padding.Smoothing (ms)
Freeze one frameHolds one encoded frame for a new requested duration.Position 0–1, duration
Jitterz' = z + amount × N(0,1)Amount, Seed
NoneLeaves the encoded latent unchanged.Use this when you want only the Energy layer.
Dimension numbering: the interface is 1-based. Dimension 1 is often the highest-variance dimension in PCA-ordered exports, but the script does not assume that this interpretation holds for every custom model.

Latent Energy — controlled instability

The Energy layer is applied after the deterministic latent operation and before decode():

encode → latent operation → optional energy → decode
ControlOptionsMeaning
Variation modeOff / Seeded / New each runOff is deterministic. Seeded is reproducible. New each run draws a new seed and reports it so the exact variation can later be reproduced in Seeded mode.
Energy typeJitter / Smooth driftJitter is frame-to-frame instability. Drift low-pass-filters the noise for correlated motion.
Energy amount0–1Strength of injected variation.
Smoothness0–1Drift only. Maps to an approximate correlation time from 0.05 s to 5 s.
Energy targetAll / One / High-varianceChooses which encoded dimensions receive the added energy.
Energy scaleRelative / AbsoluteRelative scales noise by each dimension's spread in the encoded input; Absolute uses latent-coordinate units directly.

Reproducibility in v1.5.1

The Jitter operation and the Energy layer now use independent random streams. One user Seed still reproduces the complete result, but the two noise sources are no longer identical. A Fresh run reports its seed; entering that seed in Seeded mode reproduces the same Energy variation.

Parameters & advanced settings

Main dialog

ParameterDefaultDescription
RAVE modelbreak.tsSeven preset filenames or Custom.
ActionProcess soundForward processing, latent transform, or diagnostics.
Gain (dB)0Applied after normalisation.
Normalizepeaknone = model output untouched; peak = peak 0.99; rms = target −20 dBFS with peak guard.
Draw visualisationyesDraws the processing figure after audio actions.
Play resultyesPlays the newly imported Sound.

Advanced settings

SettingOptions
Input shapeAuto-detect / BCT / B1T mono / CT
Output channelsSame as model / Force mono / Force stereo
SR handlingResample to model rate (recommended) / Use as-is (v1.2 behaviour)
Output rateBack to input rate / Keep model rate
DeviceCPU / Auto (CUDA if available) / CUDA
Show latentEnable/disable latent export for the figure
Output prefixDefault rave_out

Visualisation

The v1.5.1 figure is designed to show what the algorithm actually did rather than a generic neural-audio illustration.

PanelContent
TitleModel, source, operation/energy information or normalisation summary, and detected model rate.
SpectrogramsInput and output on the same dB scale for meaningful comparison.
WaveformsInput in grey and processed output in blue, on a shared amplitude/time display.
True latent trajectoryWhen encode() is available: up to the first eight dimensions through time, plus a z1×z2 path. In Latent transform mode, the encoded input is shown in grey and the transformed/energised latent in colour.
Feature proxy fallbackIf no usable latent export exists, the panel is explicitly labelled Input feature proxy (NOT the model latent space).
SummaryMethods, latent size, device, sample-rate handling, level processing, channel/rate result and latent/energy description.

Diagnostics

Diagnostics require no selected Sound. They load the model and report what the TorchScript export actually exposes, including:

This capability-first design is intentional: custom TorchScript/nn~ exports are not assumed to expose the same methods or metadata as the bundled RAVE models.

Technical notes

Praat selected Sound → temporary WAV + key=value parameter file Python load model → introspect metadata → optional resample to model rate → forward() OR encode() → transform → energy → decode() → optional resample to requested output rate → normalise once → gain → 32-bit float WAV + manifest + latent CSV Praat import result → Info report → visualisation → cleanup → optional playback

Normalisation

Sample rate

If the model declares a sample rate and it differs from the Sound, the default path resamples the input to the model rate with a rational windowed-sinc resampler. The result can then be returned to the original input rate or kept at the model rate. If the export does not declare a rate, the script uses the input rate as-is rather than guessing.

Output format

The Python worker writes a 32-bit IEEE float WAV. This avoids the unnecessary PCM16 quantisation used in the older runner and allows Normalize = none to remain genuinely unscaled.

Troubleshooting