IRCAM Multichannel to Binaural — User Guide
Renders a multichannel Praat Sound to binaural stereo through IRCAM Spat5 virtual speakers. Input channels are treated as loudspeaker feeds at the selected layout positions, filtered with the chosen HRTF/HRIR data, and summed to the left and right ears with automatic headroom protection.
What this does
IRCAM Multichannel to Binaural converts a loudspeaker-oriented multichannel Sound into a two-channel headphone rendering. Praat exports a temporary multichannel WAV; the Python bridge calls spat5.virtualspeakers~; Spat5 applies the selected loudspeaker layout, HRTF, ITD scale, and optional room model; the resulting binaural WAV is analysed and returned to Praat.
Quick start
- Install IRCAM Spat5 and locate the folder containing
spat5.virtualspeakers~(orspat5.virtualspeakers~.exeon Windows). - Place
spat_binaural_bridge.pyin the Praat AudioTools script folder. - Select exactly one multichannel Sound in Praat.
- Run
IRCAM_Multichannel_to_Binaural.praat. - Choose Auto-detect from channel count unless you have a specific reason to select a layout manually.
- Choose an HRTF preset and ITD value.
- Leave Headroom on Auto safe render for normal use.
- Optionally mute the LFE channel, normalize the accepted binaural render to the requested output peak, draw the visualization, or play the result.
Signal flow & safe render
Each binaural ear is a sum of several HRIR-filtered input channels. Even when no individual input channel clips, the summed binaural result can exceed full scale. The tool therefore does not rely on a single fixed pre-gain.
In Auto safe render, the first attenuation is estimated from source peak, number of active channels, and additional HRTF headroom. Every rendered output is then measured. If PCM clipping is detected, that render is rejected and Spat5 is called again from the original source with another 6 dB of attenuation, up to five attempts.
Supported layouts & channel order
| Channels | Layout | Expected order |
|---|---|---|
| 1 | Mono | C |
| 2 | 2.0 | L, R |
| 4 | 4.0 | FL, FR, BL, BR |
| 5 | 5.0 | L, R, C, Ls, Rs |
| 6 | 5.1 | L, R, C, LFE, Ls, Rs |
| 7 | 7.0 | Spat5 7.0 order |
| 8 | 7.1 | L, R, C, LFE, Ls, Rs, Lss, Rss |
| 10 | 7.1.2 | 7.1 + TpFL, TpFR |
| 12 | 7.1.4 | 7.1 + TpFL, TpFR, TpBL, TpBR |
| 24 | 22.2 | Spat5 22.2 / NHK order |
Auto-detect maps the Sound's channel count to one of these layouts. If a layout is chosen manually, the script verifies that its required channel count matches the Sound before rendering.
HRTF, ITD & room
KEMAR presets
| Preset | HRTF | Room |
|---|---|---|
| KEMAR / neutral | KEMAR | none |
| KEMAR / hall | KEMAR | hall |
| KEMAR / studio | KEMAR | studio |
Custom SOFA
Choose SOFA custom / none to supply a SOFA dataset name through Sofa_file. In this mode the separate Room preset becomes active and may be set to none, hall, livingroom, or studio.
ITD percent
ITD_percent scales the interaural time-difference component used by the Spat5 render. 100% is the normal reference value; lower values reduce temporal left/right separation and higher values exaggerate it.
Headroom & output level
Auto safe render
This is the recommended mode. The tool estimates a conservative initial input attenuation, renders the binaural signal, measures the result, and automatically tries again with an additional 6 dB attenuation if clipping is found.
Manual pre-gain
Manual mode performs one render using Manual_pre_gain_dB. The gain is still limited when necessary so that the temporary input WAV itself does not exceed approximately −1 dBFS. The rendered binaural file is still analysed for clipping.
Normalise output
When Normalise_output is enabled, a clean accepted render is scaled so that its final peak reaches Output_peak_dBFS. This happens only after the Spat5 render has passed the clipping check.
If Spat5 returns a floating-point WAV whose samples exceed ±1, the values have not been hard-clipped. The script preserves the render and scales it into a safe playback range.
LFE handling
Mute_LFE is available for layouts whose LFE position is known in the script: 5.1, 7.1, 7.1.2, and 7.1.4. In these layouts channel 4 is treated as LFE and can be set to zero before binaural rendering.
The option is ignored for layouts that do not have a known LFE channel.
Channel-order test Sound
Enable Create_channel_order_test_Sound to create a diagnostic multichannel Sound instead of rendering the selected source.
The generated test contains one approximately 0.8-second pink-noise burst per channel, one channel at a time, at about −20 dBFS peak. Burst k starts near second k−1.
- Choose the layout you want to verify.
- Create the test Sound.
- Run the binaural tool again on that test Sound using KEMAR / neutral and Auto safe render.
- Listen to the bursts in sequence and confirm that each appears from the expected loudspeaker direction.
Controls
| Control | Default | Role |
|---|---|---|
| Tools_folder | Spat5 tools path | Folder containing spat5.virtualspeakers~. |
| Layout_preset | Auto-detect | Speaker layout used to interpret the input channels. |
| Hrtf_preset | KEMAR / neutral | Selects KEMAR or custom SOFA rendering. |
| Sofa_file | kemar | SOFA dataset name used when Custom SOFA is selected. |
| Itd_percent | 100 | Scales interaural time difference. |
| Room_preset | none | Room used with Custom SOFA; KEMAR presets define their own room setting. |
| Headroom | Auto safe render | Automatic iterative protection or one-pass manual gain. |
| Manual_pre_gain_dB | −6 dB | Requested pre-gain in Manual mode. |
| Normalise_output | On | Normalize an accepted clean render to the requested output peak. |
| Output_peak_dBFS | −1 dBFS | Target peak when output normalization is enabled. |
| Mute_LFE | Off | Mute channel 4 in supported LFE layouts. |
| Create_channel_order_test_Sound | Off | Create a diagnostic one-channel-at-a-time test instead of rendering. |
| Draw_visualization | On | Draw the AudioTools diagnostic figure. |
| Play_result | On | Play a clean accepted result after processing. |
Visualization
The Picture-window figure shows both the audio result and the level-management process.
- A — Input: channel 1 and, when present, channel 2 at their true source level.
- B — Binaural output: left and right ear waveforms, with ±1 full-scale guides and final peak.
- C — Safe render: one column per render attempt; clipped attempts are distinguished from the accepted clean render, together with the input attenuation used.
- D — Level ledger: source peak, accepted input attenuation, render peak, normalization gain, output peak, and net source-to-output gain.
- Summary strip: layout, assumed channel order, HRTF, ITD, room, headroom mode, output format, duration, and descriptive L/R correlation.
Render diagnostics
The Python bridge analyses the binaural WAV after each render. It reports:
- output format, sample rate, frame count, and stereo channel count;
- peak and RMS level for each ear;
- crest factor;
- number of full-scale PCM samples;
- longest consecutive full-scale run;
- for floating-point output, samples whose absolute value exceeds 1;
- L/R zero-lag correlation.
Output behavior
A successful render creates a stereo Sound named:
Temporary WAV, log, and statistics files are deleted after a clean accepted render. If the external render fails, the log is preserved and printed in the Info window. If repeated clipping prevents acceptance, the final clipped render may be loaded for inspection but is explicitly reported as unaccepted.
The Info window also prints a level ledger containing the source peak, accepted input attenuation, number of render attempts, Spat5 render peak, normalization gain, final output peak, net gain, and L/R correlation.
Troubleshooting
Spat5 tool not found
Set Tools_folder to the directory that actually contains spat5.virtualspeakers~ or spat5.virtualspeakers~.exe.
Python not found
The script probes common Python commands and requires a working Python 3 installation on the system PATH. NumPy is optional; when available it speeds up output analysis.
Layout rejected
The selected layout must match the Sound's channel count. Auto-detect supports 1, 2, 4, 5, 6, 7, 8, 10, 12, and 24 channels.
Spatial positions sound wrong
The likely cause is channel-order mismatch. Generate and render the channel-order test Sound to verify the layout one channel at a time.
Render clips
Use Auto safe render. The tool will progressively lower the input and repeat the Spat5 render. A clipped PCM render is never accepted as clean.
Output is quieter or louder than expected
Inspect the Level ledger. The final level depends on the attenuation required for a clean render and on the optional output normalization target.
L/R correlation is high
This is not automatically an error. Frontal or strongly correlated material can produce high binaural correlation. Judge the spatial result together with the source layout and listening position.
Dependencies
- Praat
- IRCAM Spat5 command-line tools, including
spat5.virtualspeakers~ - Python 3
- NumPy is optional and used only to accelerate file analysis when available