Acoustic DNA Resonator — Differentiable FDN

Extracts a sound's spectral envelope, frequency-dependent decay profile, and modal peaks, then trains a differentiable Feedback Delay Network (FDN) to inherit selected parts of that "acoustic DNA." The original sound is convolved with the trained resonator, optionally combined with a short velvet-noise early-reflection field, and returned as a self-derived resonant transformation.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 0.2.9 (2026) License: MIT License Repo: https://github.com/ShaiCohen-ops/Praat-plugin_AudioTools
Contents:

What this does

The script analyses the selected Sound and derives a robust mono reference for training: ordinary multichannel material uses the channel mean, while severe phase cancellation falls back to the strongest channel. From that reference it measures an STFT-based spectral envelope, 24 log-spaced decay bands, and informational modal peaks. A differentiable FDN is then optimized against the decay profile and a controllable amount of the source's spectral shape.

The important idea: the wet signal is the source passed through the resonator. If the resonator itself copies the full spectral tilt of the source, that tilt is effectively applied twice. Spectral_dna therefore controls how strongly the resonator inherits the source's spectral envelope: 0 keeps the resonator spectrally flat, 1 applies the full analysed source shape, and the default 0.4 keeps a partial spectral imprint while the frequency-dependent decay profile remains an important part of the learned DNA.

Key features:

Performance: the default descriptor path does not render a time-domain impulse response on every training epoch. Decay is predicted from the model parameters and the spectral loss is evaluated on a sparse log-frequency grid, so training is much lighter than the retained legacy STFT path. Exact run time remains hardware-dependent.

Quick start

  1. In Praat, select exactly one Sound object.
  2. Optionally select one TextGrid as well if you want to export first-tier, non-empty intervals as metadata.
  3. Run script… → AcousticDNAResonator.praat.
  4. Choose Custom or one of the three named presets.
  5. Set the desired FDN size, IR duration, epochs, dry/wet balance, early-reflection window, spectral-DNA amount, normalization, output channels, and seed.
  6. Click OK. Praat locates a Python interpreter with the required packages, exports the source, launches the Python engine, imports the result, and optionally draws the analysis/report figure.
  7. The result appears as originalname_dnares; it is played automatically when Play_result is enabled.
Python requirements: numpy, scipy, soundfile, and torch. The Praat front end first tries its configured virtual environment and then OS-appropriate fallback commands. If none can import all four packages, the error lists the candidates that were tried. A typical installation command is pip install numpy scipy soundfile torch.
TextGrid note: event export is metadata only. It does not change the analysis target, training, or sound of the result.

Presets

The form opens on Custom. The three named presets overwrite only FDN size, IR duration, epochs, and dry/wet. They do not overwrite Early_reflections_ms, Spectral_dna, normalization, output channels, seed, or the display/play switches.

PresetFDN SizeIR DurationEpochsDry/WetCharacter
Custom (default)164.0 s8000.35Starting values for manual design.
Bright shimmer chamber102.0 s6000.55Compact, short and relatively wet. The name is a musical label; the preset does not add a dedicated high-shelf or forced bright EQ.
Dark long decay246.0 s10000.45Larger network and long tail. The label does not impose an extra dark spectral tilt; tonal shaping still follows the learned DNA settings.
Subtle enhancement122.5 s5000.18Mostly dry, with a lighter resonant contribution.

How it works

1. Acoustic-DNA analysis

The engine measures a spectral representation of the source, estimates post-peak decay constants in 24 logarithmic frequency bands, and detects prominent modal peaks for reporting. Decay fitting starts at each band's strongest frame and follows the measurable tail rather than fitting the attack and sustain as if they were decay.

2. Differentiable Feedback Delay Network

FDN structure

The resonator contains N parallel delay lines with fixed prime-derived lengths mᵢ, an orthogonal Householder feedback matrix A, per-line one-pole damping, and trainable input/output gains.

Householder matrix: A = I − 2vvᵀ, with v normalized to unit length.

Per-line delay/damping term: Dᵢ(z) = z⁻ᵐⁱ · g₀,ᵢ(1 − aᵢ) / (1 − aᵢz⁻¹).

FDN transfer: H(z) = g_outᵀ (I − D(z)A)⁻¹ D(z) g_in.

Sherman–Morrison structure
With Householder feedback, I − D·A = (I − D) + 2Dvvᵀ: a diagonal matrix plus a rank-1 update. The transfer can therefore be evaluated without constructing and LU-solving a full complex N × N system at every frequency. This is the key structural optimization used by the engine.

3. Default descriptor training

The Praat front end requests the descriptor loss. Training runs at the source's native sample rate and combines:

The best finite checkpoint is restored after training. The older stft_decay objective remains in the backend as a compatibility path; the Praat front end retries it only if an older engine does not accept the descriptor mode.

4. Spectral DNA without double tilt

The spectral target is mean-centered in log magnitude and multiplied by Spectral_dna. Because the source itself is later convolved with the learned resonator, a value of 1 can strongly reinforce the source's existing spectral tilt. The default 0.4 deliberately asks the resonator to inherit only part of that shape.

QC interpretation: spectral_match_mae is still measured against the full source spectral shape. Therefore lowering Spectral_dna can make that QC number worse even when the audible result is intentionally better balanced. This is expected, not a training failure.

5. Exact render and early/late split

After optimization, the final impulse response is generated by an exact causal Householder-FDN recursion at the source sample rate. If Early_reflections_ms is greater than zero, the beginning of each FDN impulse response is faded in and a seeded velvet-noise early-reflection pattern is added over that window. The early pattern uses randomized ±1 taps on a regular grid, an exponential envelope, predelay, and gentle high-frequency damping; its energy is scaled to the energy removed from the original early FDN segment.

The original source is then convolved with the resulting impulse response. Dry/wet uses an equal-power cosine/sine crossfade, not a simple linear interpolation.

Parameters

ParameterRange / choicesDefaultDescription
PresetCustom + 3 named presetsCustomNamed presets overwrite only FDN size, IR duration, epochs, and dry/wet.
Fdn_size4–3216Number of FDN delay lines. Larger networks provide more internal paths and more decorrelation possibilities.
Ir_duration0.25–15 s4.0 sLength of the final trained FDN impulse response used for convolution.
Epochs10–5000800Optimization iterations. The best finite checkpoint is used.
Dry_wet0–10.35Equal-power dry/wet control: 0 = dry only, 1 = wet only.
Early_reflections_ms0–250 ms40 msLength of the synthetic velvet-noise early-reflection field. 0 disables the split and leaves a pure FDN impulse response.
Spectral_dna0–10.4How much of the source's mean-centered spectral shape is imposed on the resonator. 0 = spectrally flat target; 1 = full analysed spectral shape.
Normalize_modenone / peak / rmsrmspeak targets 0.98 peak amplitude; rms matches the true multichannel input RMS with a safety cap on extreme gain.
Out_channelsmatch input / mono / stereo / 4 / 6 / 8match inputOutput channel count. Match-input mode is capped at 8 channels.
Seedinteger42Controls model initialization, decorrelated output taps, and seeded early-reflection patterns.
Export_textgrid_events_metadataon / offoffExports non-empty intervals from tier 1 of a simultaneously selected interval TextGrid. Metadata only; no sonic effect.
Draw_visualizationon / offonDraws the analysis/training summary figure in Praat.
Play_resulton / offonPlays the resulting Sound after processing.
Multichannel excitation: when the output channel count exactly matches the input channel count and is greater than one, each output channel is convolved with its own input channel through a decorrelated tap of the same trained FDN. If the counts do not match, all output channels are excited from the robust mono analysis reference instead.

Output & visualisation

The processed Sound is imported into Praat as originalname_dnares. Because the render is convolution with an IR of length Ir_duration, the output extends beyond the source by approximately the IR length minus one sample. The Python interchange WAV is written as 32-bit float before Praat re-imports it, avoiding hidden PCM-16 clipping in the handoff.

Visualisation panels

When Draw_visualization is enabled, the script draws:
  • Input waveform — the selected source Sound.
  • Rendered output spectrogram — channel 1 of the processed result, up to 20 kHz or Nyquist.
  • Training loss — up to 400 sampled loss values, including the final training region.
  • Decay DNA comparison — input-band decay values versus the exact trained/rendered IR on the same frequency bands.
  • Summary panel — FDN size, epochs, best epoch, loss values, decay estimate, spectral MAE, decay log-MAE, duration, channel count, dry/wet, normalization, RMS, analysis reference, and delay lengths.
Analysis reference: normal multichannel material is analysed from the arithmetic mean. If that mean exhibits severe phase cancellation relative to the strongest channel, the engine switches the analysis/render reference to the strongest channel and reports that choice.

Applications

Self-derived resonance

Transform a recording through a resonator whose frequency-dependent decay and spectral tendency are learned from the recording itself. This works well for voice, instrumental tones, environmental recordings, and other material where the source's own resonant signature is compositionally meaningful.

Controlled spectral inheritance

Use Spectral_dna as a compositional control between a more neutral FDN spectrum and a stronger imprint of the source envelope. Lower settings are useful when the wet result becomes too dark or over-emphasizes the source's existing tilt.

Hybrid early/late spatial character

The early-reflection window adds a short, diffuse attack before the learned FDN tail. This can make percussive, plucked, and speech-like material feel less sparse at the onset while retaining the FDN as the late resonant field.

Multichannel resonant processing

With Out_channels = match input, stereo or multichannel sources retain channel-specific excitation and dry content while different output taps decorrelate the resonant field.

Workflow: Voice → long self-resonant tail

Start: Dark long decay preset.
Then adjust: Spectral_dna downward if formant/tilt reinforcement becomes too strong; keep Early_reflections_ms around 40 ms for a more immediate onset, or set it to 0 for the pure learned FDN response.
Result: a long resonant tail shaped by the voice's measured decay profile without requiring the resonator to copy the full source spectral tilt.

Workflow: Stereo field recording → preserved-channel resonance

Start: Subtle enhancement, Out_channels = match input.
Result: left and right channels excite their own decorrelated taps of the same trained FDN while the original stereo channels remain separate in the dry path.

Troubleshooting

Python or dependencies not found: install numpy scipy soundfile torch into a Python environment the script can reach, or edit the configured pyCandidate1$ path near the top of the Praat script.

Engine version mismatch: the Praat front end prints the exact Python engine path it loaded. Update that exact acoustic_dna_resonator.py file so the front end and engine are from the same release.

Wet result is too dark: lower Spectral_dna. The source already contributes its own spectral tilt before the resonator response is applied.

Attack feels sparse or delayed: increase Early_reflections_ms from 0 toward the default 40 ms. The FDN itself cannot respond before its shortest delay and takes time to build a dense late field.

Spectral MAE becomes worse after lowering Spectral_dna: expected. That QC metric still compares the trained transfer against the full source spectral shape, while the model is intentionally targeting only a fraction of it.

TextGrid seems to do nothing: expected. Exported intervals are metadata only and do not drive the current sonic model.

Long run: reduce Epochs first. On the descriptor path, IR duration mainly affects final rendering rather than the per-epoch training objective.