IRCAM Multichannel to Binaural — User Guide

Renders a multichannel Praat Sound to binaural stereo through IRCAM Spat5 virtual speakers. Input channels are treated as loudspeaker feeds at the selected layout positions, filtered with the chosen HRTF/HRIR data, and summed to the left and right ears with automatic headroom protection.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel License: MIT License Repo: Praat AudioTools
Contents:

What this does

IRCAM Multichannel to Binaural converts a loudspeaker-oriented multichannel Sound into a two-channel headphone rendering. Praat exports a temporary multichannel WAV; the Python bridge calls spat5.virtualspeakers~; Spat5 applies the selected loudspeaker layout, HRTF, ITD scale, and optional room model; the resulting binaural WAV is analysed and returned to Praat.

This is a virtual-loudspeaker renderer, not a generic stereo downmix. Each input channel must correspond to the loudspeaker position expected by the selected layout. Correct channel order is therefore essential.

Quick start

  1. Install IRCAM Spat5 and locate the folder containing spat5.virtualspeakers~ (or spat5.virtualspeakers~.exe on Windows).
  2. Place spat_binaural_bridge.py in the Praat AudioTools script folder.
  3. Select exactly one multichannel Sound in Praat.
  4. Run IRCAM_Multichannel_to_Binaural.praat.
  5. Choose Auto-detect from channel count unless you have a specific reason to select a layout manually.
  6. Choose an HRTF preset and ITD value.
  7. Leave Headroom on Auto safe render for normal use.
  8. Optionally mute the LFE channel, normalize the accepted binaural render to the requested output peak, draw the visualization, or play the result.
For an unfamiliar multichannel file, use Create channel-order test Sound first. It is the quickest way to verify that the channel order you expect matches the Spat5 layout.

Signal flow & safe render

MULTICHANNEL PRAAT SOUND ↓ optional LFE mute ↓ input gain / automatic headroom ↓ temporary multichannel WAV ↓ spat5.virtualspeakers~ layout + HRTF + ITD + room ↓ BINAURAL WAV ↓ peak / RMS / full-scale / clipping analysis ↓ clipped? YES → discard render reduce input by 6 dB render again NO → accept ↓ optional output normalization ↓ STEREO PRAAT SOUND

Each binaural ear is a sum of several HRIR-filtered input channels. Even when no individual input channel clips, the summed binaural result can exceed full scale. The tool therefore does not rely on a single fixed pre-gain.

In Auto safe render, the first attenuation is estimated from source peak, number of active channels, and additional HRTF headroom. Every rendered output is then measured. If PCM clipping is detected, that render is rejected and Spat5 is called again from the original source with another 6 dB of attenuation, up to five attempts.

Clipping is prevented by re-rendering, not repaired afterwards. Output normalization is applied only after a clean render has been accepted.

Supported layouts & channel order

ChannelsLayoutExpected order
1MonoC
22.0L, R
44.0FL, FR, BL, BR
55.0L, R, C, Ls, Rs
65.1L, R, C, LFE, Ls, Rs
77.0Spat5 7.0 order
87.1L, R, C, LFE, Ls, Rs, Lss, Rss
107.1.27.1 + TpFL, TpFR
127.1.47.1 + TpFL, TpFR, TpBL, TpBR
2422.2Spat5 22.2 / NHK order

Auto-detect maps the Sound's channel count to one of these layouts. If a layout is chosen manually, the script verifies that its required channel count matches the Sound before rendering.

Channel count is not enough to prove channel order. Praat knows that a Sound has, for example, eight channels; it does not know whether those channels were assembled in the order expected by 7.1. Use the channel-order test when in doubt.

HRTF, ITD & room

KEMAR presets

PresetHRTFRoom
KEMAR / neutralKEMARnone
KEMAR / hallKEMARhall
KEMAR / studioKEMARstudio

Custom SOFA

Choose SOFA custom / none to supply a SOFA dataset name through Sofa_file. In this mode the separate Room preset becomes active and may be set to none, hall, livingroom, or studio.

ITD percent

ITD_percent scales the interaural time-difference component used by the Spat5 render. 100% is the normal reference value; lower values reduce temporal left/right separation and higher values exaggerate it.

Headroom & output level

Auto safe render

This is the recommended mode. The tool estimates a conservative initial input attenuation, renders the binaural signal, measures the result, and automatically tries again with an additional 6 dB attenuation if clipping is found.

Manual pre-gain

Manual mode performs one render using Manual_pre_gain_dB. The gain is still limited when necessary so that the temporary input WAV itself does not exceed approximately −1 dBFS. The rendered binaural file is still analysed for clipping.

If a manual render clips, the result is not accepted as a clean output. Use Auto safe render or lower the manual pre-gain.

Normalise output

When Normalise_output is enabled, a clean accepted render is scaled so that its final peak reaches Output_peak_dBFS. This happens only after the Spat5 render has passed the clipping check.

If Spat5 returns a floating-point WAV whose samples exceed ±1, the values have not been hard-clipped. The script preserves the render and scales it into a safe playback range.

LFE handling

Mute_LFE is available for layouts whose LFE position is known in the script: 5.1, 7.1, 7.1.2, and 7.1.4. In these layouts channel 4 is treated as LFE and can be set to zero before binaural rendering.

The option is ignored for layouts that do not have a known LFE channel.

Channel-order test Sound

Enable Create_channel_order_test_Sound to create a diagnostic multichannel Sound instead of rendering the selected source.

The generated test contains one approximately 0.8-second pink-noise burst per channel, one channel at a time, at about −20 dBFS peak. Burst k starts near second k−1.

  1. Choose the layout you want to verify.
  2. Create the test Sound.
  3. Run the binaural tool again on that test Sound using KEMAR / neutral and Auto safe render.
  4. Listen to the bursts in sequence and confirm that each appears from the expected loudspeaker direction.
If single-channel bursts are clean but a dense multichannel source distorts, the problem is likely summing/headroom. If an individual burst is already wrong or distorted, inspect layout, channel order, HRTF, sample rate, and Spat5 configuration.

Controls

ControlDefaultRole
Tools_folderSpat5 tools pathFolder containing spat5.virtualspeakers~.
Layout_presetAuto-detectSpeaker layout used to interpret the input channels.
Hrtf_presetKEMAR / neutralSelects KEMAR or custom SOFA rendering.
Sofa_filekemarSOFA dataset name used when Custom SOFA is selected.
Itd_percent100Scales interaural time difference.
Room_presetnoneRoom used with Custom SOFA; KEMAR presets define their own room setting.
HeadroomAuto safe renderAutomatic iterative protection or one-pass manual gain.
Manual_pre_gain_dB−6 dBRequested pre-gain in Manual mode.
Normalise_outputOnNormalize an accepted clean render to the requested output peak.
Output_peak_dBFS−1 dBFSTarget peak when output normalization is enabled.
Mute_LFEOffMute channel 4 in supported LFE layouts.
Create_channel_order_test_SoundOffCreate a diagnostic one-channel-at-a-time test instead of rendering.
Draw_visualizationOnDraw the AudioTools diagnostic figure.
Play_resultOnPlay a clean accepted result after processing.

Visualization

The Picture-window figure shows both the audio result and the level-management process.

Render diagnostics

The Python bridge analyses the binaural WAV after each render. It reports:

L/R correlation is descriptive only. A frontal binaural source can legitimately be highly correlated, while diffuse or lateral material can be much less correlated. Correlation is not used as a pass/fail measure of binaural quality.

Output behavior

A successful render creates a stereo Sound named:

originalName_binaural

Temporary WAV, log, and statistics files are deleted after a clean accepted render. If the external render fails, the log is preserved and printed in the Info window. If repeated clipping prevents acceptance, the final clipped render may be loaded for inspection but is explicitly reported as unaccepted.

The Info window also prints a level ledger containing the source peak, accepted input attenuation, number of render attempts, Spat5 render peak, normalization gain, final output peak, net gain, and L/R correlation.

Troubleshooting

Spat5 tool not found

Set Tools_folder to the directory that actually contains spat5.virtualspeakers~ or spat5.virtualspeakers~.exe.

Python not found

The script probes common Python commands and requires a working Python 3 installation on the system PATH. NumPy is optional; when available it speeds up output analysis.

Layout rejected

The selected layout must match the Sound's channel count. Auto-detect supports 1, 2, 4, 5, 6, 7, 8, 10, 12, and 24 channels.

Spatial positions sound wrong

The likely cause is channel-order mismatch. Generate and render the channel-order test Sound to verify the layout one channel at a time.

Render clips

Use Auto safe render. The tool will progressively lower the input and repeat the Spat5 render. A clipped PCM render is never accepted as clean.

Output is quieter or louder than expected

Inspect the Level ledger. The final level depends on the attenuation required for a clean render and on the optional output normalization target.

L/R correlation is high

This is not automatically an error. Frontal or strongly correlated material can produce high binaural correlation. Judge the spatial result together with the source layout and listening position.

Dependencies