Algorithmic Metallic Synthesis — User Guide

Generates layered metallic textures from regularly triggered, exponentially decaying FM ring kernels.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 0.4 (2026) License: MIT License Repo: https://github.com/ShaiCohen-ops/Praat-plugin_AudioTools
Contents:

What this does

Algorithmic Metallic Synthesis creates a new sound from scratch; no input Sound object is required. Each synthesis voice consists of a regular strike train convolved with an exponentially decaying FM ring kernel. The kernel supplies an inharmonic carrier, a decaying modulation index, and an amplitude decay, while the strike train determines when that resonance is excited.

Because the complete ring kernel is reproduced at every strike, resonance tails can overlap naturally when strikes occur faster than the decay. Multiple voices and layers are summed, then the result receives a global edge fade, optional stereo rendering, optional target-peak normalization, visualization, and playback.

Quick start

  1. In Praat, run Run script… → Algorithmic Metallic Synthesis.praat. No sound selection is needed.
  2. Choose a Preset, or keep Custom to use the settings below the preset menu.
  3. Set the synthesis and output controls as needed.
  4. Click OK. The generated object is named metallic_<preset>.

For a clear first example, try Temple Bell for separated, long-decay resonances or Gamelan for a denser multi-voice texture.

Synthesis mechanism

Strike train × ring kernel

For each voice, the script creates a unit pulse train on exact sample positions. The nominal strike rate is converted to an integer period in samples, so the realized rate is sample_rate / period_samples. With randomization enabled, each voice also receives a random initial offset within one period.

The pulse train is convolved with the voice's ring kernel:

h(t) = A · exp(-t / τ) ·
       sin(2πfc t + β · exp(-t / τM) · sin(2πfm t))

τ  = realized resonance decay
τM = 0.60 · τ
fc = carrier frequency after detuning and anti-alias adjustment
fm = layer modulation rate
β  = modulation depth

The kernel normally lasts five decay constants, with a 30 ms minimum and a cap at the requested output duration. Its modulation index decays faster than its amplitude, so each ring tends to move from a more strongly modulated onset toward a simpler oscillation as it decays.

Overlap and level compensation

Strike tails are allowed to overlap. To prevent strike density from becoming an unintended gain control, each voice is attenuated when its average overlap estimate strike_rate × decay exceeds one. Voice amplitude is also scaled by the number of layers and by the square root of the number of voices in that layer.

Parameters

ParameterDefaultBehavior
PresetCustomSelects Custom or one of seven named configurations. Presets override only the fields listed in the Presets section.
Duration_s3Duration of the generated output. Convolution tails extending beyond this boundary are cropped.
Sample_rate_Hz44100Output sample rate. Values below 8000 Hz are rejected. The safe synthesis ceiling is 45% of the sample rate.
Base_frequency_Hz200Reference frequency from which each mode constructs its carrier lattice.
Number_of_voices5Requested voices per layer. Clamped to 1–12. Dense Shimmer doubles this count; Sparse Bells uses approximately half.
Number_of_layers3Number of synthesis layers, clamped to 1–8.
Modulation_rate_Hz0.5Base FM modulation rate. With randomization, each layer varies from 0.7× to 1.3× this value.
Modulation_depth3.0Base phase-modulation index β. Each synthesis mode applies its own multiplier, and randomization may vary it further.
Resonance_decay_s0.1Per-strike exponential amplitude decay constant τ. With randomization, each layer varies from 0.8× to 1.2× this value.
Synthesis_modeStandard MetallicSelects the carrier lattice, voice count, strike-rate pattern, amplitude scaling, and modulation-depth multiplier.
Fade_time_s0.5Global linear fade-in and fade-out applied after synthesis. Negative values become 0; values above half the duration are limited to half the duration.
Spatial_modeMonoMono, Stereo Wide, or equal-power Rotating output.
Randomize_parametersYesControls stochastic variation of layer/voice parameters and strike offsets. It does not change the regular pulse-train structure within a voice.
Random_seed00 uses an unpredictable random stream; a positive integer makes randomized synthesis reproducible.
Normalize_outputYesIf the output is nonzero, scales the final global peak to 0.9.
Draw_visualizationYesDraws the measured/model visualization described below.
Play_resultYesPlays the final Sound after generation.

Presets

Preset selection occurs before validation. Fields not listed for a preset retain the values currently entered in the form. In particular, presets do not change Sample_rate_Hz, Random_seed, Normalize_output, Draw_visualization, or Play_result.

PresetFields overridden
Temple BellDuration 5 s; base 180 Hz; 6 voices; 2 layers; modulation rate 0.3 Hz; depth 4.0; decay 0.8 s; Sparse Bells; Rotating; fade 1 s.
GamelanDuration 4 s; base 250 Hz; 8 voices; 3 layers; modulation rate 0.8 Hz; depth 2.5; decay 0.3 s; Dense Shimmer; Stereo Wide.
Industrial ClangDuration 3 s; base 120 Hz; 6 voices; 4 layers; modulation rate 2.0 Hz; depth 5.0; decay 0.05 s; Rhythmic Clang; Stereo Wide.
Wind ChimesDuration 6 s; base 800 Hz; 10 voices; 2 layers; modulation rate 0.2 Hz; depth 1.5; decay 0.4 s; Sparse Bells; Rotating; Randomize_parameters = Yes.
Gong WashDuration 8 s; base 80 Hz; 4 voices; 3 layers; modulation rate 0.1 Hz; depth 6.0; decay 1.5 s; Standard Metallic; Rotating; fade 2 s.
Prepared PianoDuration 4 s; base 150 Hz; 6 voices; 3 layers; modulation rate 1.5 Hz; depth 3.5; decay 0.15 s; Chaotic Resonance; Stereo Wide.
Steel DrumDuration 3 s; base 300 Hz; 5 voices; 2 layers; modulation rate 1.0 Hz; depth 2.0; decay 0.2 s; Rhythmic Clang; Mono.

Synthesis modes

ModeVoice / carrier structureStrike structure
Standard MetallicRequested voice count. Carriers begin at the third multiple of the base and rise with voice and layer; 2% progressive detuning. Uses the full modulation depth.Voice-dependent regular rates beginning at 3 Hz.
Dense ShimmerTwice the requested voices per layer, higher carrier placement, 1.5% progressive detuning, and 0.7× modulation depth.Faster voice-dependent regular rates beginning at 4 Hz.
Sparse Bellsmax(2, floor(voices/2)) voices per layer, wider layer spacing, 1% progressive detuning, and 1.3× modulation depth.Lower voice-dependent regular rates beginning at 2 Hz.
Rhythmic ClangRequested voice count, carriers based on integer multiples of the base with layer expansion, 2.5% progressive detuning, and 0.8× modulation depth.Regular rates of 2 × voice Hz before sample-grid quantization.
Chaotic ResonanceWith randomization enabled, frequency, detuning, modulation depth, and strike rate are drawn from bounded random ranges. With randomization disabled, the mode uses a fixed irregular carrier/rate lattice.Randomized 2–8 Hz nominal rates when randomization is on; fixed irregular rates when it is off.

The Chaotic Resonance label describes the mode's irregular behavior. Its control law is stochastic when randomization is enabled and a fixed irregular lattice when disabled, rather than a named deterministic-chaos equation.

Spatial modes

ModeProcessingChannels
MonoThe synthesized mono signal is retained directly.1
Stereo WideLeft = original / √2. Right = the same signal delayed by 3 ms / √2. The right-channel delay is cropped at the fixed output boundary.2
RotatingComplementary sine/cosine equal-power gains move the mono source across stereo at 0.3 Hz.2

Both stereo modes are created from the synthesized mono signal; there is no pre-existing spatial image to preserve.

Randomization & seed

With Randomize_parameters = Yes, the script randomizes layer modulation rate and decay, voice frequency and modulation depth, and the initial strike offset. Chaotic Resonance additionally randomizes its carrier construction, detuning, modulation depth, and strike rate.

A positive Random_seed reproduces these randomized choices. Seed 0 uses an unpredictable stream. With Randomize_parameters = No, the synthesis parameters and strike offsets are deterministic, so the audible result does not depend on the seed.

Visualization

When Draw_visualization is enabled, the script draws a model/measurement view from the actual values used by the DSP.

A — Measured output waveform

Shows the final output waveform on a symmetric amplitude scale. For stereo output, the channel with the higher RMS is displayed; the channels are not folded to mono.

B — Actual strike-resonance field

Plots carrier frequency against time. Rust vertical ticks mark realized strikes; pale horizontal lines show each strike's actual ring-kernel duration. For large event counts, the drawing is subsampled for legibility while the synthesis itself remains unchanged.

C — Model → measurement

Paints a measured spectrogram of the same representative output channel and overlays horizontal guides at the actual realized carrier frequencies. When many voices are present, carrier guides are subsampled for legibility.

The lower mechanism strip summarizes the pulse-train/convolution model. The final dashboard reports realized carrier range, mean decay, strike count, anti-alias adjustments, skipped voices, output peak/RMS, spatial mode, and seed.

Output & safeguards

Compositional context

Metallic timbre synthesis has a long computer-music lineage built around inharmonic frequency relationships, time-varying spectra, and decaying resonances. This tool approaches that territory through a specific algorithmic design: regularly triggered pulse trains convolved with decaying FM ring kernels, allowing the same resonance model to range from isolated bell-like strikes to dense overlapping metallic fields.