Acoustic DNA Resonator — Differentiable FDN
Extracts a sound's spectral envelope, frequency-dependent decay profile, and modal peaks, then trains a differentiable Feedback Delay Network (FDN) to inherit selected parts of that "acoustic DNA." The original sound is convolved with the trained resonator, optionally combined with a short velvet-noise early-reflection field, and returned as a self-derived resonant transformation.
What this does
The script analyses the selected Sound and derives a robust mono reference for training: ordinary multichannel material uses the channel mean, while severe phase cancellation falls back to the strongest channel. From that reference it measures an STFT-based spectral envelope, 24 log-spaced decay bands, and informational modal peaks. A differentiable FDN is then optimized against the decay profile and a controllable amount of the source's spectral shape.
0 keeps the resonator spectrally flat, 1 applies the full analysed source shape, and the default 0.4 keeps a partial spectral imprint while the frequency-dependent decay profile remains an important part of the learned DNA.
Key features:
- Descriptor-based training by default — directly optimizes decay and sparse spectral descriptors at the source's native sample rate.
- Differentiable FDN — fixed prime delays, a trainable Householder feedback vector, one-pole frequency-dependent damping, and trainable input/output gains.
- Exact final FDN render — the trained impulse response is generated by a causal time-domain recursion rather than by a periodically sampled IFFT approximation.
- Early/late split — an optional velvet-noise early-reflection pattern improves the first tens of milliseconds before the learned FDN late field develops.
- Self-excitation — the selected sound excites the resonator learned from that same sound.
- Multichannel output — up to 8 decorrelated output taps; when input and output channel counts match, each channel is excited by its own source channel.
- Built-in QC — the report compares input-vs-trained-IR decay behaviour and reports spectral and decay match errors.
descriptor path does not render a time-domain impulse response on every training epoch. Decay is predicted from the model parameters and the spectral loss is evaluated on a sparse log-frequency grid, so training is much lighter than the retained legacy STFT path. Exact run time remains hardware-dependent.
Quick start
- In Praat, select exactly one Sound object.
- Optionally select one TextGrid as well if you want to export first-tier, non-empty intervals as metadata.
- Run script… →
AcousticDNAResonator.praat. - Choose Custom or one of the three named presets.
- Set the desired FDN size, IR duration, epochs, dry/wet balance, early-reflection window, spectral-DNA amount, normalization, output channels, and seed.
- Click OK. Praat locates a Python interpreter with the required packages, exports the source, launches the Python engine, imports the result, and optionally draws the analysis/report figure.
- The result appears as
originalname_dnares; it is played automatically when Play_result is enabled.
numpy, scipy, soundfile, and torch. The Praat front end first tries its configured virtual environment and then OS-appropriate fallback commands. If none can import all four packages, the error lists the candidates that were tried. A typical installation command is pip install numpy scipy soundfile torch.
Presets
The form opens on Custom. The three named presets overwrite only FDN size, IR duration, epochs, and dry/wet. They do not overwrite Early_reflections_ms, Spectral_dna, normalization, output channels, seed, or the display/play switches.
| Preset | FDN Size | IR Duration | Epochs | Dry/Wet | Character |
|---|---|---|---|---|---|
| Custom (default) | 16 | 4.0 s | 800 | 0.35 | Starting values for manual design. |
| Bright shimmer chamber | 10 | 2.0 s | 600 | 0.55 | Compact, short and relatively wet. The name is a musical label; the preset does not add a dedicated high-shelf or forced bright EQ. |
| Dark long decay | 24 | 6.0 s | 1000 | 0.45 | Larger network and long tail. The label does not impose an extra dark spectral tilt; tonal shaping still follows the learned DNA settings. |
| Subtle enhancement | 12 | 2.5 s | 500 | 0.18 | Mostly dry, with a lighter resonant contribution. |
How it works
1. Acoustic-DNA analysis
The engine measures a spectral representation of the source, estimates post-peak decay constants in 24 logarithmic frequency bands, and detects prominent modal peaks for reporting. Decay fitting starts at each band's strongest frame and follows the measurable tail rather than fitting the attack and sustain as if they were decay.
2. Differentiable Feedback Delay Network
FDN structure
The resonator contains N parallel delay lines with fixed prime-derived lengths mᵢ, an orthogonal Householder feedback matrix A, per-line one-pole damping, and trainable input/output gains.
Householder matrix: A = I − 2vvᵀ, with v normalized to unit length.
Per-line delay/damping term: Dᵢ(z) = z⁻ᵐⁱ · g₀,ᵢ(1 − aᵢ) / (1 − aᵢz⁻¹).
FDN transfer: H(z) = g_outᵀ (I − D(z)A)⁻¹ D(z) g_in.
With Householder feedback, I − D·A = (I − D) + 2Dvvᵀ: a diagonal matrix plus a rank-1 update. The transfer can therefore be evaluated without constructing and LU-solving a full complex N × N system at every frequency. This is the key structural optimization used by the engine.
3. Default descriptor training
The Praat front end requests the descriptor loss. Training runs at the source's native sample rate and combines:
- Closed-form decay matching — per-line decay constants are derived from attenuation per delay, then combined into a model of the band envelope and compared with the analysed source decay.
- Sparse spectral-shape matching — the FDN transfer is sampled at 6 log-spaced points inside each of up to 28 log-frequency bands rather than on a dense FFT grid.
- Stability guard — a small penalty acts only near the feedback-gain stability boundary.
The best finite checkpoint is restored after training. The older stft_decay objective remains in the backend as a compatibility path; the Praat front end retries it only if an older engine does not accept the descriptor mode.
4. Spectral DNA without double tilt
The spectral target is mean-centered in log magnitude and multiplied by Spectral_dna. Because the source itself is later convolved with the learned resonator, a value of 1 can strongly reinforce the source's existing spectral tilt. The default 0.4 deliberately asks the resonator to inherit only part of that shape.
spectral_match_mae is still measured against the full source spectral shape. Therefore lowering Spectral_dna can make that QC number worse even when the audible result is intentionally better balanced. This is expected, not a training failure.
5. Exact render and early/late split
After optimization, the final impulse response is generated by an exact causal Householder-FDN recursion at the source sample rate. If Early_reflections_ms is greater than zero, the beginning of each FDN impulse response is faded in and a seeded velvet-noise early-reflection pattern is added over that window. The early pattern uses randomized ±1 taps on a regular grid, an exponential envelope, predelay, and gentle high-frequency damping; its energy is scaled to the energy removed from the original early FDN segment.
The original source is then convolved with the resulting impulse response. Dry/wet uses an equal-power cosine/sine crossfade, not a simple linear interpolation.
Parameters
| Parameter | Range / choices | Default | Description |
|---|---|---|---|
| Preset | Custom + 3 named presets | Custom | Named presets overwrite only FDN size, IR duration, epochs, and dry/wet. |
| Fdn_size | 4–32 | 16 | Number of FDN delay lines. Larger networks provide more internal paths and more decorrelation possibilities. |
| Ir_duration | 0.25–15 s | 4.0 s | Length of the final trained FDN impulse response used for convolution. |
| Epochs | 10–5000 | 800 | Optimization iterations. The best finite checkpoint is used. |
| Dry_wet | 0–1 | 0.35 | Equal-power dry/wet control: 0 = dry only, 1 = wet only. |
| Early_reflections_ms | 0–250 ms | 40 ms | Length of the synthetic velvet-noise early-reflection field. 0 disables the split and leaves a pure FDN impulse response. |
| Spectral_dna | 0–1 | 0.4 | How much of the source's mean-centered spectral shape is imposed on the resonator. 0 = spectrally flat target; 1 = full analysed spectral shape. |
| Normalize_mode | none / peak / rms | rms | peak targets 0.98 peak amplitude; rms matches the true multichannel input RMS with a safety cap on extreme gain. |
| Out_channels | match input / mono / stereo / 4 / 6 / 8 | match input | Output channel count. Match-input mode is capped at 8 channels. |
| Seed | integer | 42 | Controls model initialization, decorrelated output taps, and seeded early-reflection patterns. |
| Export_textgrid_events_metadata | on / off | off | Exports non-empty intervals from tier 1 of a simultaneously selected interval TextGrid. Metadata only; no sonic effect. |
| Draw_visualization | on / off | on | Draws the analysis/training summary figure in Praat. |
| Play_result | on / off | on | Plays the resulting Sound after processing. |
Output & visualisation
The processed Sound is imported into Praat as originalname_dnares. Because the render is convolution with an IR of length Ir_duration, the output extends beyond the source by approximately the IR length minus one sample. The Python interchange WAV is written as 32-bit float before Praat re-imports it, avoiding hidden PCM-16 clipping in the handoff.
Visualisation panels
- Input waveform — the selected source Sound.
- Rendered output spectrogram — channel 1 of the processed result, up to 20 kHz or Nyquist.
- Training loss — up to 400 sampled loss values, including the final training region.
- Decay DNA comparison — input-band decay values versus the exact trained/rendered IR on the same frequency bands.
- Summary panel — FDN size, epochs, best epoch, loss values, decay estimate, spectral MAE, decay log-MAE, duration, channel count, dry/wet, normalization, RMS, analysis reference, and delay lengths.
Applications
Self-derived resonance
Transform a recording through a resonator whose frequency-dependent decay and spectral tendency are learned from the recording itself. This works well for voice, instrumental tones, environmental recordings, and other material where the source's own resonant signature is compositionally meaningful.
Controlled spectral inheritance
Use Spectral_dna as a compositional control between a more neutral FDN spectrum and a stronger imprint of the source envelope. Lower settings are useful when the wet result becomes too dark or over-emphasizes the source's existing tilt.
Hybrid early/late spatial character
The early-reflection window adds a short, diffuse attack before the learned FDN tail. This can make percussive, plucked, and speech-like material feel less sparse at the onset while retaining the FDN as the late resonant field.
Multichannel resonant processing
With Out_channels = match input, stereo or multichannel sources retain channel-specific excitation and dry content while different output taps decorrelate the resonant field.
Workflow: Voice → long self-resonant tail
Start: Dark long decay preset.
Then adjust: Spectral_dna downward if formant/tilt reinforcement becomes too strong; keep Early_reflections_ms around 40 ms for a more immediate onset, or set it to 0 for the pure learned FDN response.
Result: a long resonant tail shaped by the voice's measured decay profile without requiring the resonator to copy the full source spectral tilt.
Workflow: Stereo field recording → preserved-channel resonance
Start: Subtle enhancement, Out_channels = match input.
Result: left and right channels excite their own decorrelated taps of the same trained FDN while the original stereo channels remain separate in the dry path.
Troubleshooting
numpy scipy soundfile torch into a Python environment the script can reach, or edit the configured pyCandidate1$ path near the top of the Praat script.Engine version mismatch: the Praat front end prints the exact Python engine path it loaded. Update that exact
acoustic_dna_resonator.py file so the front end and engine are from the same release.Wet result is too dark: lower Spectral_dna. The source already contributes its own spectral tilt before the resonator response is applied.
Attack feels sparse or delayed: increase Early_reflections_ms from 0 toward the default 40 ms. The FDN itself cannot respond before its shortest delay and takes time to build a dense late field.
Spectral MAE becomes worse after lowering Spectral_dna: expected. That QC metric still compares the trained transfer against the full source spectral shape, while the model is intentionally targeting only a fraction of it.
TextGrid seems to do nothing: expected. Exported intervals are metadata only and do not drive the current sonic model.
Long run: reduce Epochs first. On the descriptor path, IR duration mainly affects final rendering rather than the per-epoch training objective.