Algorithmic Metallic Synthesis — User Guide
Generates layered metallic textures from regularly triggered, exponentially decaying FM ring kernels.
What this does
Algorithmic Metallic Synthesis creates a new sound from scratch; no input Sound object is required. Each synthesis voice consists of a regular strike train convolved with an exponentially decaying FM ring kernel. The kernel supplies an inharmonic carrier, a decaying modulation index, and an amplitude decay, while the strike train determines when that resonance is excited.
Because the complete ring kernel is reproduced at every strike, resonance tails can overlap naturally when strikes occur faster than the decay. Multiple voices and layers are summed, then the result receives a global edge fade, optional stereo rendering, optional target-peak normalization, visualization, and playback.
Quick start
- In Praat, run Run script… →
Algorithmic Metallic Synthesis.praat. No sound selection is needed. - Choose a Preset, or keep Custom to use the settings below the preset menu.
- Set the synthesis and output controls as needed.
- Click OK. The generated object is named
metallic_<preset>.
For a clear first example, try Temple Bell for separated, long-decay resonances or Gamelan for a denser multi-voice texture.
Synthesis mechanism
Strike train × ring kernel
For each voice, the script creates a unit pulse train on exact sample positions. The nominal strike rate is converted to an integer period in samples, so the realized rate is sample_rate / period_samples. With randomization enabled, each voice also receives a random initial offset within one period.
The pulse train is convolved with the voice's ring kernel:
h(t) = A · exp(-t / τ) ·
sin(2πfc t + β · exp(-t / τM) · sin(2πfm t))
τ = realized resonance decay
τM = 0.60 · τ
fc = carrier frequency after detuning and anti-alias adjustment
fm = layer modulation rate
β = modulation depth
The kernel normally lasts five decay constants, with a 30 ms minimum and a cap at the requested output duration. Its modulation index decays faster than its amplitude, so each ring tends to move from a more strongly modulated onset toward a simpler oscillation as it decays.
Overlap and level compensation
Strike tails are allowed to overlap. To prevent strike density from becoming an unintended gain control, each voice is attenuated when its average overlap estimate strike_rate × decay exceeds one. Voice amplitude is also scaled by the number of layers and by the square root of the number of voices in that layer.
Parameters
| Parameter | Default | Behavior |
|---|---|---|
| Preset | Custom | Selects Custom or one of seven named configurations. Presets override only the fields listed in the Presets section. |
| Duration_s | 3 | Duration of the generated output. Convolution tails extending beyond this boundary are cropped. |
| Sample_rate_Hz | 44100 | Output sample rate. Values below 8000 Hz are rejected. The safe synthesis ceiling is 45% of the sample rate. |
| Base_frequency_Hz | 200 | Reference frequency from which each mode constructs its carrier lattice. |
| Number_of_voices | 5 | Requested voices per layer. Clamped to 1–12. Dense Shimmer doubles this count; Sparse Bells uses approximately half. |
| Number_of_layers | 3 | Number of synthesis layers, clamped to 1–8. |
| Modulation_rate_Hz | 0.5 | Base FM modulation rate. With randomization, each layer varies from 0.7× to 1.3× this value. |
| Modulation_depth | 3.0 | Base phase-modulation index β. Each synthesis mode applies its own multiplier, and randomization may vary it further. |
| Resonance_decay_s | 0.1 | Per-strike exponential amplitude decay constant τ. With randomization, each layer varies from 0.8× to 1.2× this value. |
| Synthesis_mode | Standard Metallic | Selects the carrier lattice, voice count, strike-rate pattern, amplitude scaling, and modulation-depth multiplier. |
| Fade_time_s | 0.5 | Global linear fade-in and fade-out applied after synthesis. Negative values become 0; values above half the duration are limited to half the duration. |
| Spatial_mode | Mono | Mono, Stereo Wide, or equal-power Rotating output. |
| Randomize_parameters | Yes | Controls stochastic variation of layer/voice parameters and strike offsets. It does not change the regular pulse-train structure within a voice. |
| Random_seed | 0 | 0 uses an unpredictable random stream; a positive integer makes randomized synthesis reproducible. |
| Normalize_output | Yes | If the output is nonzero, scales the final global peak to 0.9. |
| Draw_visualization | Yes | Draws the measured/model visualization described below. |
| Play_result | Yes | Plays the final Sound after generation. |
Presets
Preset selection occurs before validation. Fields not listed for a preset retain the values currently entered in the form. In particular, presets do not change Sample_rate_Hz, Random_seed, Normalize_output, Draw_visualization, or Play_result.
| Preset | Fields overridden |
|---|---|
| Temple Bell | Duration 5 s; base 180 Hz; 6 voices; 2 layers; modulation rate 0.3 Hz; depth 4.0; decay 0.8 s; Sparse Bells; Rotating; fade 1 s. |
| Gamelan | Duration 4 s; base 250 Hz; 8 voices; 3 layers; modulation rate 0.8 Hz; depth 2.5; decay 0.3 s; Dense Shimmer; Stereo Wide. |
| Industrial Clang | Duration 3 s; base 120 Hz; 6 voices; 4 layers; modulation rate 2.0 Hz; depth 5.0; decay 0.05 s; Rhythmic Clang; Stereo Wide. |
| Wind Chimes | Duration 6 s; base 800 Hz; 10 voices; 2 layers; modulation rate 0.2 Hz; depth 1.5; decay 0.4 s; Sparse Bells; Rotating; Randomize_parameters = Yes. |
| Gong Wash | Duration 8 s; base 80 Hz; 4 voices; 3 layers; modulation rate 0.1 Hz; depth 6.0; decay 1.5 s; Standard Metallic; Rotating; fade 2 s. |
| Prepared Piano | Duration 4 s; base 150 Hz; 6 voices; 3 layers; modulation rate 1.5 Hz; depth 3.5; decay 0.15 s; Chaotic Resonance; Stereo Wide. |
| Steel Drum | Duration 3 s; base 300 Hz; 5 voices; 2 layers; modulation rate 1.0 Hz; depth 2.0; decay 0.2 s; Rhythmic Clang; Mono. |
Synthesis modes
| Mode | Voice / carrier structure | Strike structure |
|---|---|---|
| Standard Metallic | Requested voice count. Carriers begin at the third multiple of the base and rise with voice and layer; 2% progressive detuning. Uses the full modulation depth. | Voice-dependent regular rates beginning at 3 Hz. |
| Dense Shimmer | Twice the requested voices per layer, higher carrier placement, 1.5% progressive detuning, and 0.7× modulation depth. | Faster voice-dependent regular rates beginning at 4 Hz. |
| Sparse Bells | max(2, floor(voices/2)) voices per layer, wider layer spacing, 1% progressive detuning, and 1.3× modulation depth. | Lower voice-dependent regular rates beginning at 2 Hz. |
| Rhythmic Clang | Requested voice count, carriers based on integer multiples of the base with layer expansion, 2.5% progressive detuning, and 0.8× modulation depth. | Regular rates of 2 × voice Hz before sample-grid quantization. |
| Chaotic Resonance | With randomization enabled, frequency, detuning, modulation depth, and strike rate are drawn from bounded random ranges. With randomization disabled, the mode uses a fixed irregular carrier/rate lattice. | Randomized 2–8 Hz nominal rates when randomization is on; fixed irregular rates when it is off. |
The Chaotic Resonance label describes the mode's irregular behavior. Its control law is stochastic when randomization is enabled and a fixed irregular lattice when disabled, rather than a named deterministic-chaos equation.
Spatial modes
| Mode | Processing | Channels |
|---|---|---|
| Mono | The synthesized mono signal is retained directly. | 1 |
| Stereo Wide | Left = original / √2. Right = the same signal delayed by 3 ms / √2. The right-channel delay is cropped at the fixed output boundary. | 2 |
| Rotating | Complementary sine/cosine equal-power gains move the mono source across stereo at 0.3 Hz. | 2 |
Both stereo modes are created from the synthesized mono signal; there is no pre-existing spatial image to preserve.
Randomization & seed
With Randomize_parameters = Yes, the script randomizes layer modulation rate and decay, voice frequency and modulation depth, and the initial strike offset. Chaotic Resonance additionally randomizes its carrier construction, detuning, modulation depth, and strike rate.
A positive Random_seed reproduces these randomized choices. Seed 0 uses an unpredictable stream. With Randomize_parameters = No, the synthesis parameters and strike offsets are deterministic, so the audible result does not depend on the seed.
Visualization
When Draw_visualization is enabled, the script draws a model/measurement view from the actual values used by the DSP.
A — Measured output waveform
Shows the final output waveform on a symmetric amplitude scale. For stereo output, the channel with the higher RMS is displayed; the channels are not folded to mono.
B — Actual strike-resonance field
Plots carrier frequency against time. Rust vertical ticks mark realized strikes; pale horizontal lines show each strike's actual ring-kernel duration. For large event counts, the drawing is subsampled for legibility while the synthesis itself remains unchanged.
C — Model → measurement
Paints a measured spectrogram of the same representative output channel and overlays horizontal guides at the actual realized carrier frequencies. When many voices are present, carrier guides are subsampled for legibility.
The lower mechanism strip summarizes the pulse-train/convolution model. The final dashboard reports realized carrier range, mean decay, strike count, anti-alias adjustments, skipped voices, output peak/RMS, spatial mode, and seed.
Output & safeguards
- Object name:
metallic_<preset>; Custom producesmetallic_Custom. - Duration: exactly the requested output duration after any preset override. Convolution tails beyond the endpoint are cropped.
- Sample rate: the selected Sample_rate_Hz.
- Normalization: when enabled and the signal is nonzero, Scale peak: 0.9 performs target peak normalization. When disabled, the script reports a warning if the peak exceeds 0.99.
- Anti-alias guard: carriers are kept below
0.45 × sample_ratewith additional margin for the estimated FM sideband spread. A carrier may be moved downward; a voice is skipped if no safe carrier remains. - Workload guards: synthesis is rejected above 160 estimated oscillator voices or above the script's 100,000-strike workload estimate.
Compositional context
Metallic timbre synthesis has a long computer-music lineage built around inharmonic frequency relationships, time-varying spectra, and decaying resonances. This tool approaches that territory through a specific algorithmic design: regularly triggered pulse trains convolved with decaying FM ring kernels, allowing the same resonance model to range from isolated bell-like strikes to dense overlapping metallic fields.