Beltrami Inspired Spectral Melter — User Guide
An edge-aware time-frequency diffusion instrument: it turns a spectrogram into a dB terrain, diffuses that terrain anisotropically, exaggerates the resulting spectral shape, and resynthesizes it by overlap-add.
What this does
Beltrami Inspired Spectral Melter treats the source spectrogram as a two-dimensional terrain whose axes are time and frequency. It converts spectral energy to dB, measures local gradients, diffuses the terrain more freely in smooth areas than across strong ridges, exaggerates the diffused spectral shape, then resynthesizes a new Sound by overlap-add.
The result is not a conventional filter or reverb. It is a time-varying spectral morphology: energy can smear through neighbouring time frames and frequency bins while prominent ridges can be protected to different degrees.
Beltrami and anisotropic diffusion
The Laplace–Beltrami operator generalizes the ordinary Laplacian from flat Euclidean space to curved Riemannian manifolds. In broad terms, Laplacian diffusion smooths a field by letting values flow toward their neighbours.
Anisotropic diffusion makes that flow direction- or content-dependent. Here the script computes a local gradient magnitude g and uses a conductance term of the form:
Small gradients receive conductance near 1 and diffuse readily. Strong ridges receive lower conductance and are protected. Time and frequency have separate diffusion strengths.
Quick start
- Select exactly one Sound.
- Run
Beltrami_Inspired_Spectral_Melter.praat. - Choose one of the seven named presets or Custom.
- Choose mono or stereo output with Create_stereo.
- For stereo, use Stereo_phase_offset: 0 gives matching phase fields; 1 gives the widest independent phase difference.
- Choose a Speed mode.
- Use Edit_details only when you need analysis, diffusion, strength, wet/dry or seed control.
Presets
Diffusion and mix
| Preset | Window / step | Iterations | Time / freq diffusion | Strength | Wet |
|---|---|---|---|---|---|
| Custom | 40 / 10 ms | 6 | 0.15 / 0.12 | 3.0 | 0.85 |
| Shimmer Haze | 40 / 10 ms | 8 | 0.20 / 0.05 | 2.5 | 0.80 |
| Deep Terrain | 60 / 15 ms | 12 | 0.20 / 0.18 | 4.0 | 0.90 |
| Edge Freeze | 30 / 8 ms | 5 | 0.10 / 0.08 | 2.0 | 0.70 |
| Fog of War | 50 / 12 ms | 10 | 0.12 / 0.22 | 4.0 | 0.88 |
| Formant Cloud | 35 / 8 ms | 7 | 0.18 / 0.06 | 3.0 | 0.75 |
| Transient Glass | 25 / 6 ms | 6 | 0.22 / 0.04 | 3.5 | 0.82 |
| Void Chasm | 100 / 40 ms | 18 | 0.24 / 0.20 | 8.0 | 1.00 |
Analysis and edge protection
| Preset | Max Hz | Freq res. | Floor dB | Ridge sensitivity | Edge preservation |
|---|---|---|---|---|---|
| Custom | 6000 | 100 Hz | -80 | 1.8 | 1.0 |
| Shimmer Haze | 6000 | 80 Hz | -80 | 2.5 | 1.2 |
| Deep Terrain | 5000 | 100 Hz | -70 | 1.2 | 0.8 |
| Edge Freeze | 7000 | 70 Hz | -80 | 3.5 | 2.0 |
| Fog of War | 5000 | 120 Hz | -75 | 1.5 | 0.9 |
| Formant Cloud | 5000 | 50 Hz | -80 | 2.0 | 1.5 |
| Transient Glass | 8000 | 80 Hz | -80 | 4.0 | 2.5 |
| Void Chasm | 12000 | 150 Hz | -80 | 2.5 | 1.5 |
Reevaluate edges each iteration remains off unless changed in Edit details. Create_stereo, Stereo_phase_offset, Speed mode and Random seed also remain user choices. Max frequency is reduced when necessary to stay below the working Nyquist frequency by at least one requested frequency-resolution step.
Main controls
| Control | Default | Meaning |
|---|---|---|
| Preset | Custom | Loads analysis, diffusion, effect-strength and wet/dry defaults. |
| Create_stereo | On | On = stereo randomized-phase resynthesis; off = mono phase-preserving resynthesis. |
| Stereo_phase_offset | 1.0 | 0…1. Controls the additional right-channel phase difference relative to the shared base phase field. |
| Speed_mode | Full quality | Full = original rate; Balanced = process at up to 22.05 kHz; Fast = process at up to 11.025 kHz. |
| Edit_details | Off | Opens all analysis/diffusion parameters after the preset has been loaded. |
| Draw_visualization | On | Draws the measured terrain, diffusion law, spectral slice and final output. |
| Play_result | On | Plays the final Sound. |
Edit details
| Control | Meaning |
|---|---|
| Effect strength | Expands deviations from each frame's mean linear amplitude. 1 = use the diffused shape as-is; larger values exaggerate peaks/valleys before a positivity clamp. |
| Wet/dry mix | Linear 0…1 blend. |
| Window size / Time step | Gaussian spectrogram analysis geometry. The requested window is rounded up to a power-of-two sample length. |
| Max frequency | Upper analysis frequency; automatically limited by the working Nyquist frequency. |
| Frequency resolution | Requested spectrogram frequency sampling. |
| Dynamic floor | Lower dB floor used when converting the spectrogram to a log-energy terrain; must be below 0 dB. |
| Diffusion iterations | Number of repeated diffusion passes. |
| Time diffusion / Frequency diffusion | Separate finite-difference step sizes. Internally each is capped at 0.24. |
| Ridge sensitivity / Edge preservation | Together set κ = ridge_sensitivity / edge_preservation. Higher edge preservation lowers κ and protects ridges more strongly. |
| Reevaluate edges each iteration | Recomputes the gradient field from the evolving terrain each pass. |
| Random seed | 0 = a new stereo phase realization; positive integer = repeatable stereo phase realization. |
Processing pipeline
- Convert multichannel input to mono.
- Optionally downsample according to Speed mode.
- Build a Gaussian spectrogram and convert it to a dB terrain.
- Measure time/frequency gradient magnitude.
- Run edge-aware anisotropic diffusion for the requested iterations.
- Convert the diffused dB terrain back to linear amplitude.
- Expand spectral deviations around each frame's arithmetic mean by
Effect strength; clamp negative amplitudes to zero. - Resynthesize by 75%-overlap Hann overlap-add, using the target spectral magnitudes.
- Mix wet/dry, trim to source duration, apply 10 ms / 20 ms final fades, restore original sample rate when needed, and target-normalize to 0.95.
Mono / stereo phase model
The magnitude transformation is shared, but the resynthesis phase law depends on output mode.
| Mode | Phase behaviour |
|---|---|
| Mono | Preserves the source-frame complex phase while replacing the magnitude with the diffused target magnitude. |
| Stereo | Uses a seeded random phase field. Left uses the shared base field; right adds an independent phase-difference term scaled by Stereo_phase_offset. Offset 0 approaches dual mono; offset 1 is widest. |
Channels, speed and level
- Input topology: multichannel input is converted to mono before analysis and processing.
- Output topology: mono when Create_stereo is off; stereo when it is on.
- Duration: trimmed to the source duration.
- Sample rate: final output is restored to the original rate after Balanced/Fast processing when the original rate was higher than the target.
- Wet scaling: stereo wet channels are joint-scaled to 0.99; mono wet is target-scaled to 0.99.
- Final level: the complete output is always target-normalized to 0.95.
- 0% wet is not a transparent bypass: the source has already been monoized and may have been resampled; final fades and 0.95 target normalization still occur.
Output name: <source>_BeltramiInspired_<preset>, with _stereo appended in stereo mode.
Visualization
- A — Spectral terrain: source log-energy terrain → diffused terrain → actual resynthesis target.
- B — Diffusion law: conductance curve and finite-difference flow, including κ and time/frequency step sizes.
- C — Measured spectral slice: source versus target at the frame where the terrain changed most.
- D — Final output: mono or stereo waveform on a shared amplitude scale.
Historical / technical / compositional context
Eugenio Beltrami introduced the Laplace–Beltrami equation in the 1860s as a generalization of Laplace's equation to curved surfaces. The name of this tool points toward that geometric idea: a field whose local geometry determines how smoothing or diffusion should proceed.
The implemented algorithm, however, is technically closer to Pietro Perona and Jitendra Malik's 1990 anisotropic diffusion: smoothing is reduced at strong gradients so edges can survive while flatter regions flow. The optional per-iteration edge reevaluation makes that relationship especially direct.
Compositionally, the script treats a spectrogram as a malleable terrain. Formants, partial ridges and transient edges become topographic features that may resist or permit diffusion. Time diffusion creates memory and smearing; frequency diffusion melts vertical spectral structure; contrast expansion turns the smoothed terrain into a new, often exaggerated resynthesis target. This is best understood as a spectral-morphology instrument, not as a physical simulation of a manifold.
Further reading
- Encyclopedia of Mathematics — Laplace–Beltrami equation — definition and historical note on Beltrami's 1864–1865 work.
- P. Perona and J. Malik, “Scale-Space and Edge Detection Using Anisotropic Diffusion,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 12(7), 629–639, 1990. IEEE Xplore.