Generative Sound System — User Guide

Six distinct generative synthesis engines in one interface: stochastic harmonic drift, Poisson granular synthesis, logistic-map chaotic FM, spectral-band morphing, probabilistic rhythmic pulses, and subtractive noise with evolving band gains.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 0.4.1 (2026) License: MIT License Repo: Praat AudioTools
Contents:

What this does

Generative Sound System is not one synthesis algorithm with several presets. It contains six genuinely different sound-generation processes that share the same duration, base-frequency, layer, evolution, spatial, seed, fade, level, and visualization framework.

Architecture:
chosen generative process → independent layers/events → layer-level spatialization → one global edge fade → optional final target peak normalization

The synthesis engine and spatial mode are independent choices. Spatialization is applied at the layer level, not as a single post-mix stereo effect.

Quick start

  1. Run Generative_Sound_System.praat. No input Sound is required.
  2. Choose one of the six Synthesis modes.
  3. Set Duration, Sample rate, Base frequency, Number of layers, and Evolution rate.
  4. Choose Mono, Layer Spread, Rotating Layers, or Micro-Delay Stereo.
  5. Open Edit generative details when you want mode-specific stochastic controls.
  6. Leave Normalize output on for target peak normalization to 0.90, or turn it off to retain the raw generated level.

Shared controls

ControlRange / defaultMeaning
Durationdefault 10 s; max 120 sExact output duration. Events reaching the end are shortened to remain inside it.
Sample ratedefault 44.1 kHz; 8–192 kHzFinal output sample rate and basis for the practical frequency-headroom check.
Base frequencydefault 110 HzStarting reference used differently by each synthesis mode.
Number of layers1–16; default 4Number of parallel harmonic voices, grain populations, carriers, spectral bands, pulse layers, or noise bands.
Evolution rate>0 to 20; default 0.5Active in every mode, but its exact meaning is mode-dependent.
Fade timedefault 1 sOne global linear fade at both outer edges, capped at half the total duration.
Normalize outputon by defaultApplies one final target peak normalization to 0.90 after spatialization and fade.

How Evolution rate changes by mode

ModeRole of Evolution rate
Harmonic DriftControls the speed of each stochastic frequency-drift process and therefore its control rate.
Granular CloudRaises the expected Poisson grain density in every layer.
Chaotic FMRaises the logistic-map control rate and moves the logistic parameter slightly toward stronger mixing.
Spectral MorphSets the frequency of the coherent spectral-focus motion across the band bank.
Rhythmic PulseSets the base pulse-event rate for each layer.
Subtractive NoiseControls the speed of each independent stochastic band-gain process and therefore its control rate.

Synthesis modes

1. Harmonic Drift

Layer n has nominal frequency effectiveBase × n. Each layer receives an independent bounded stochastic drift control. The control is generated at a low control rate, resampled to the audio sample rate, clamped to [-1,1] in v0.4.1, and mapped to instantaneous frequency.

Audio phase is then integrated sample by sample from that instantaneous-frequency trajectory before taking the sine. This is genuine frequency drift, not a fixed oscillator with a nested phase expression.

instantaneous frequency = nominal × (1 + driftDepth × control) layer amplitude = 1 / sqrt(layer number)

The stochastic controller is an Ornstein–Uhlenbeck-like bounded process implemented through correlated Gaussian updates followed by tanh. The v0.4.1 post-resample clamp guarantees that the final audio-rate controller stays inside the requested drift-depth range.

2. Granular Cloud

This mode uses genuine stochastic grains. Each layer has its own homogeneous Poisson onset process:

inter-onset interval ~ Exponential(layer density)

Layer centers rise linearly from the effective Base frequency using 1 + 0.30 × (layer-1). Each grain receives independent octave-domain pitch jitter, a random duration between 25 and 110 ms, random phase, and a full Hann envelope.

Grain density depends on layer number, Granular density scale, and Evolution rate, and is capped at 120 events/s per layer. Events can overlap. A common grain-amplitude factor compensates approximately for expected overlap.

Rendering is performed one layer at a time in 0.5-second local chunks. A grain crossing a chunk boundary retains its original onset, phase, frequency, duration, and Hann envelope, so chunking does not restart the grain.

3. Chaotic FM

Each layer uses a logistic map as its control source:

x[n+1] = r × x[n] × (1 - x[n])

The map runs at a mode-dependent control rate and is sinc-resampled to the final audio rate. In v0.4.1 the resampled control is explicitly clamped to [0,1] before being mapped to instantaneous carrier frequency. This prevents interpolation overshoot from exceeding the intended FM range.

instantaneous frequency = carrier × (1 + depth × (2x - 1))

As in Harmonic Drift, audio phase is then integrated sample by sample and converted to a sine. This is therefore frequency modulation by a chaotic control trajectory, not merely a sinusoidal phase formula labelled “chaos”.

4. Spectral Morph

A single Gaussian-noise source is copied into a bank of fixed Hann-pass spectral bands. Band centers span the requested number of octaves from the effective Base frequency.

A single coherent focus trajectory moves across normalized band position 0–1:

focus(t) = 0.5 + 0.5 × sin(2π × Evolution rate × t)

Each band receives a Gaussian-shaped gain according to its distance from that focus. The spectrum therefore changes because the weights of actual spectral bands change through time; this is not merely amplitude modulation before a fixed filter.

5. Rhythmic Pulse

Each layer has a nominal event rate derived from Evolution rate. Initial onset is randomized within one nominal IOI. Subsequent onset intervals receive user-controlled percentage jitter, and every scheduled onset can be omitted independently according to Pulse omission probability.

Rendered events are sine bursts with random phase and a full Hann envelope. Event duration is the smaller of 160 ms and 30% of the nominal IOI, clipped at the end of the Sound.

This creates an explicit event schedule rather than amplitude-modulating a continuous oscillator.

6. Subtractive Noise

Each layer begins as independent Gaussian noise, then passes through a fixed Hann-pass band. Band centers span the same octave-domain spectral architecture used by Spectral Morph.

Each filtered band receives its own bounded stochastic gain control. The controller is generated at a low control rate, sinc-resampled to audio rate, clamped to [-1,1] in v0.4.1, and then mapped approximately to the gain range 0.18–1.0:

gain = 0.18 + 0.82 × (0.5 + 0.5 × control)

Unlike Spectral Morph, the bands do not share one moving focus; their gain trajectories evolve independently.

Generative Details

The optional advanced page provides mode-specific controls. Controls irrelevant to the selected mode are harmless but have no effect on that run.

ControlDefaultUsed by
Random seed0All stochastic modes and random event schedules.
Harmonic drift depth3%Harmonic Drift.
Granular density scale1.0Granular Cloud.
Grain pitch jitter0.20 octGranular Cloud.
Chaotic FM depth fraction0.28Chaotic FM.
Spectral span3 octSpectral Morph and Subtractive Noise.
Noise-band width0.65 octSpectral Morph and Subtractive Noise.
Pulse timing jitter18%Rhythmic Pulse.
Pulse omission probability0.12Rhythmic Pulse.

Random seed

Seed 0 uses an unpredictable random state. A positive seed reproduces all stochastic controls and event decisions for the same settings. This includes random drift/noise controllers, grain schedules and phases, logistic-map initial states, shared or independent noise realizations as applicable, and rhythmic onset/omission decisions. After rendering, Praat's unpredictable random initialization is restored.

Spatial modes

Spatialization happens as each synthesized layer is added to the output. The layer sum is scaled by 1 / sqrt(Number of layers) before final normalization.

ModeImplementation
MonoAll layers summed into one channel.
Layer SpreadFixed equal-power positions distributed from pan 0.05 to 0.95 across the layer index.
Rotating LayersEvery layer follows a continuous equal-power pan trajectory with a unique layer phase offset. Rotation rate is 0.06 + 0.025 × Evolution rate Hz.
Micro-Delay StereoLeft receives the undelayed layer; right receives the same full-spectrum layer delayed by 0.25 ms + 0.12 ms × layer number. Both use equal scaling.
Micro-Delay Stereo is not binaural synthesis. It uses no HRTF, head-shadow model, interaural spectral filtering, or individualized ear transfer functions. The former “Binaural” label was removed because it overstated what the processing did.

Frequency headroom

The script computes a mode-dependent estimate of the highest requested fundamental/carrier/band edge. It then applies one common scale:

frequencyScale = min(1, (0.45 × sample rate) / requestedTop)

The effective Base frequency is Base frequency × frequencyScale. This preserves the relative architecture of layers instead of clipping only the highest one.

The estimate differs by mode:

Version 0.4.1 adds post-resample clamps to Harmonic Drift, Chaotic FM, and Subtractive Noise control signals, so sinc interpolation cannot push their final audio-rate controllers beyond the ranges assumed by these calculations.

Output, fade, and level

PropertyBehavior
InputNo selected Sound is required; all modes synthesize from scratch.
DurationExactly the requested duration.
Sample rateSelected 8–192 kHz rate. Some low-rate stochastic controls are internally resampled to this rate, but the final Sound remains at the requested rate.
ChannelsMono in Mono mode; stereo in all three spatial modes.
FadeOne global linear fade-in and fade-out. Actual fade = min(Fade time, 50% of Duration).
NormalizationIf enabled, one final Scale peak: 0.90 is applied after the spatial mix and global fade.
Output namegenerative_<synthesis mode> with spaces replaced by underscores.

Scale peak: 0.90 is target peak normalization, not an attenuate-only ceiling: any non-zero result is scaled so its final absolute peak becomes 0.90.

Visualization and QC

The Picture display changes its first panels according to the selected synthesis mechanism rather than forcing all six modes into the same explanatory graph.

PanelWhat it shows
A — Generative Control / Event RealizationHarmonic Drift and Chaotic FM: actual layer-1 instantaneous-frequency control. Granular Cloud and Rhythmic Pulse: actual event onset/duration by layer. Spectral Morph: coherent moving focus. Subtractive Noise: actual layer-1 stochastic band gain.
B — Frequency ArchitectureActual event frequencies for Grain/Pulse modes; otherwise nominal layer carrier or band centers, displayed on a log-frequency axis.
C — Model → MeasurementMeasured spectrogram of a representative final-output channel with mode-relevant nominal/event-frequency guides.
D — Measured OutputMeasured waveform of the same representative final-output channel after the layer-level spatial mix and final level stage.

For stereo output, the representative channel is whichever channel has the higher whole-file RMS.

The QC strip reports the active process, event count or continuous-control status, Evolution rate, seed, requested spectral top, practical safe top, applied common frequency scale, and pre/post-normalization peak/RMS.