Cluster-Based Ambient Drone Designer — User Guide

AI-driven granular synthesis: analyzes source audio for spectral stability, clusters tonal segments, and generates infinite lush ambient drones through intelligent grain recombination.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 1.1 (2026) License: MIT License Repo: https://github.com/ShaiCohen-ops/Praat-plugin_AudioTools
Contents:

What this does

This script implements neural-inspired ambient drone generation — an intelligent granular synthesis approach that analyzes source audio for spectral stability, identifies the most tonal segments through clustering, and recombines them into infinite lush ambient textures. Process involves: (1) Feature extraction: Spectral centroid, bandwidth, harmonicity, and pitch analysis across audio grains. (2) AI clustering: K-means grouping of similar spectral profiles. (3) Tonal selection: Identification of most harmonic cluster for drone foundation. (4) Generative synthesis: Randomized grain recombination with optional octave shimmer effects. Result: evolving, texturally rich ambient drones that preserve the spectral character of the source while creating entirely new sonic landscapes.

Key Features:

What are neural ambient drones? Traditional granular synthesis: manual grain selection, random recombination. Neural approach: intelligent analysis identifies "best" grains based on spectral stability and harmonic content. AI clustering: groups similar sonic characteristics automatically. Benefits: (1) Automatic quality control: System selects most musically useful segments. (2) Spectral coherence: Grains share similar acoustic properties. (3) Evolutionary textures: Gradual transitions between related sonic materials. (4) Source respect: Preserves essence while transforming completely. (5) Infinite variation: Same source yields different drones each time. Use cases: Ambient music production, soundscape design, film scoring, meditation audio, experimental composition, generative art.

Technical Implementation: (1) Preprocessing: Convert to mono, duration validation. (2) Feature extraction: Slice audio into grains (default 100ms), analyze spectral centroid (brightness), bandwidth (spectral spread), harmonicity (tonal vs noisy), pitch (fundamental frequency). (3) Normalization: Z-score standardization of features. (4) Clustering: K-means algorithm groups grains by spectral similarity. (5) Cluster selection: Choose cluster with highest average harmonicity (most tonal). (6) Granular synthesis: Random selection from tonal cluster, concatenation with overlap, optional octave shifting (50% down/200% up). (7) Layering: Multiple parallel streams mixed for density. (8) Output: Normalized, named "originalname_NeuralDrone". Processing time scales with source duration and output length.

Quick start

  1. In Praat, select exactly one Sound object.
  2. Run script… → neural_ambient_drone_designer.praat.
  3. Set output_duration_sec (e.g., 20.0 for 20-second drone).
  4. Adjust layer_density (1-5, higher = more layered texture).
  5. Enable add_octave_shimmer for harmonic enrichment.
  6. Set grain_size_ms (100ms default, smaller = more granular).
  7. Choose number_of_clusters (3 default, higher = more specific grouping).
  8. Enable play_result to audition immediately.
  9. Click OK — drone generated, named "originalname_NeuralDrone".
Quick tip: Start with tonal source material — singing, strings, pads, bells work best. Use layer_density 3-4 for rich textures. Enable octave_shimmer for harmonic complexity. Grain_size 50-150ms works well — smaller for more granular texture, larger for smoother evolution. 3-5 clusters typically sufficient. Processing shows progress: "Analyzing spectral stability..." → "Clustering textures..." → "Generating drone layers..." → "Mixing drone...". Output appears in Objects window as "originalname_NeuralDrone". Longer source audio = more grain variety. Complex drones may take 30+ seconds to generate.
Important: SOURCE SELECTION CRITICAL — script works best with tonal, sustained sources (voices, strings, pads). Noisy/percussive sources yield less musical results. Minimum source duration: At least 2× grain size (e.g., 200ms for 100ms grains). Processing time: Scales with source length and output duration — long drones from long sources can take minutes. Memory usage: Many temporary objects created during processing — close other Praat objects if memory issues. Random element: Each run produces different results from same source. Cluster selection: Automatically chooses "most tonal" cluster — may not match human perception. Grain overlap: Fixed at 50% overlap — creates smooth transitions but can cause phasing.

AI Analysis Theory

Feature Extraction

Spectral Analysis

Four-dimensional feature space:

1. SPECTRAL CENTROID (Brightness) Measure of spectral center of gravity Higher values = brighter, more high-frequency content Formula: ∑(f[i] × magn[i]) / ∑magn[i] Where f[i] = frequency bin, magn[i] = magnitude 2. SPECTRAL BANDWIDTH (Spread) Standard deviation of spectrum around centroid Higher values = wider spectral spread, more noisy Lower values = focused spectral energy, more tonal 3. HARMONICITY (Tonal vs Noisy) Harmonic-to-noise ratio from cross-correlation Higher values = more periodic, more tonal Lower values = more noisy, less structured Range: -50 dB (noisy) to +30 dB (highly tonal) 4. PITCH (Fundamental Frequency) Fundamental frequency estimation Higher values = higher pitch Zero/undefined = unpitched/noisy segments

Why These Features?

Feature selection rationale:

Together they capture:

Clustering Algorithm

K-Means Implementation

Manual K-means steps:

STEP 1: INITIALIZATION Create k random centroids in 4D feature space k = number_of_clusters (user parameter) STEP 2: E-STEP (Expectation) FOR each grain i (1 to nGrains): Calculate distance to each centroid Distance = Euclidean in 4D space: d = √[(centroid_diff)² + (bandwidth_diff)² + (harmonicity_diff)² + (pitch_diff)²] Assign grain to nearest centroid STEP 3: M-STEP (Maximization) FOR each cluster c (1 to k): Recalculate centroid as mean of assigned grains New_centroid = average(feature_vectors in cluster) STEP 4: CONVERGENCE CHECK Repeat E+M steps until assignments stable OR maximum iterations reached (10 default) RESULT: k clusters of spectrally similar grains

Why K-Means?

Advantages for audio grain clustering:

📊 Cluster Selection Logic

Selection criterion: Highest average harmonicity

Why harmonicity?

  • Best predictor of musical usefulness
  • Tonal grains create more coherent drones
  • Reduces noisy/unpitched artifacts
  • Creates foundation for harmonic development

Calculation:

For each cluster: mean_harmonicity = average(HNR values)

Select cluster with maximum mean_harmonicity

Normalization Process

Z-Score Standardization

Manual Z-score calculation:

FOR each feature column f (1 to 4): STEP 1: Calculate mean mean = (sum of all values in column) / nGrains STEP 2: Calculate standard deviation variance = sum[(value - mean)² for all grains] / nGrains sd = √variance STEP 3: Apply Z-score FOR each grain i (1 to nGrains): z = (value[i] - mean) / sd Set normalized value = z END FOR Purpose: Equalize feature scales for fair clustering Centroid, bandwidth, harmonicity, pitch have different units/ranges Z-score puts all features on same scale (mean=0, sd=1) Prevents single feature from dominating clustering

Complete Analysis Pipeline

SETUP: Select Sound object → Convert to mono → Validate duration FEATURE EXTRACTION: Slice into grains (grain_size_ms, 50% overlap) FOR each grain: Extract spectral centroid Extract spectral bandwidth Extract harmonicity (HNR) Extract pitch (F0) Create feature matrix: nGrains × 4 features NORMALIZATION: FOR each feature column: Calculate mean and standard deviation Apply Z-score: (value - mean) / sd CLUSTERING: Initialize k centroids randomly REPEAT until convergence: E-step: Assign grains to nearest centroid M-step: Recalculate centroids as cluster means SELECTION: Calculate mean harmonicity for each cluster Choose cluster with highest mean harmonicity Extract grain indices from selected cluster OUTPUT: List of "tonal grains" for synthesis phase

Granular Synthesis Engine

Grain Processing

Grain Extraction

Windowed grain creation:

Parameters: grain_size_ms = 100 (default) step_size = grain_size_ms × 0.5 = 50ms (50% overlap) Extraction: FOR each grain index i in tonal cluster: center_time = (i - 0.5) × step_size start_time = center_time - (grain_size_ms/2000) end_time = center_time + (grain_size_ms/2000) Extract sound: start_time to end_time, Hanning window Window: Hanning (cosine) reduces clicks at boundaries Overlap: 50% provides smooth transitions between grains

Octave Shimmer Effect

Pitch shifting algorithm:

Probability-based pitch shifting: FOR each grain (random selection): roll = random integer 1-10 IF roll = 1 (10% chance): Resample to 50% original rate → octave down IF roll = 10 (10% chance): Resample to 200% original rate → octave up ELSE (80% chance): No change → original pitch Resampling method: 50-point interpolation Override sampling frequency maintains timing Effect: Creates harmonic shimmer without rhythm disruption Rationale: 20% of grains shifted (10% down + 10% up) Enough for harmonic enrichment Not enough to destroy original character Random distribution avoids pattern repetition

Layer Generation

Multi-Layer Architecture

Parallel layer construction:

INPUT: layer_density = n (user parameter, 1-5 typical) FOR layer from 1 to n: STEP 1: Calculate grains needed grain_duration = grain_size_ms / 1000 overlap_factor = 0.5 (50% overlap) effective_grain_duration = grain_duration × overlap_factor grains_needed = ceil(output_duration / effective_grain_duration) STEP 2: Block-based generation WHILE grains_generated < grains_needed: Select random grains from tonal cluster Apply octave shimmer (probabilistic) Concatenate grains into blocks (20 grains/block) STEP 3: Layer assembly Concatenate all blocks → raw layer Trim to exact output_duration Store layer for mixing END FOR RESULT: n parallel layers of granular texture

Why Multiple Layers?

Sonic benefits of layering:

Mixing and Output

Final Assembly

Stereo-to-mono conversion:

IF nLayers > 1: Combine all layers to stereo (Praat built-in) Then convert stereo to mono (averaging) ELSE: Simply copy single layer Why stereo then mono? Praat's "Combine to stereo" handles multiple channels Conversion to mono ensures compatibility Maintains equal level across layers Normalization: Scale peak to 0.99 Prevents clipping while maximizing level Consistent output volume Naming: "originalname_NeuralDrone" Clear identification of processed result

Complete Synthesis Pipeline

INPUT: Tonal grain indices from analysis phase FOR each layer (1 to layer_density): grains_generated = 0 WHILE grains_generated < grains_needed: block = [] FOR 1 to 20 (or until grains_needed reached): // Random grain selection random_index = random integer from tonal indices grain = extract_grain(random_index) // Optional pitch shifting if add_octave_shimmer and random condition met: grain = resample_grain(grain, 0.5 or 2.0) block.append(grain) grains_generated += 1 // Concatenate block block_sound = concatenate(block) layer_blocks.append(block_sound) // Assemble layer from blocks raw_layer = concatenate(layer_blocks) final_layer = trim(raw_layer, output_duration) layers.append(final_layer) // Mix all layers if layers.count > 1: stereo = combine_to_stereo(layers) output = convert_to_mono(stereo) else: output = layers[1] output = normalize(output, 0.99) rename(output, original_name + "_NeuralDrone")

Visualization

When Draw_visualization is enabled, the script draws a results-oriented Praat Picture. The figure is designed to show what the analysis selected and how that selected material was actually used in synthesis, rather than repeating the parameter values already available in the form and Info window.

Reading the figure: cluster colours are categorical. The cluster selected for synthesis is always shown in green; other clusters use distinct non-green colours. The same cluster colours are reused across the cluster ribbon, feature-space plot, and cluster-composition panel so the panels can be read together.

Title and subtitle

The title identifies Cluster-Based Ambient Drone Designer v1.1. The subtitle reports the source name, active preset, number of analysis grains, active/requested cluster count, selected cluster, and number of synthesis layers.

Source waveform and cluster ribbon

The upper section links the source directly to the clustering result:

Feature Space

This panel is a two-dimensional view of the four-dimensional feature space used by k-means. Each grain is plotted by spectral centroid on the horizontal axis and HNR on the vertical axis:

If some grains have no usable harmonicity measurement, they are placed just below the lowest valid HNR region. A dotted reference line marks the valid-data boundary and the panel reports how many grains were floored there.

Important: this plot is a projection for inspection, not the complete clustering space. K-means still operates on all four normalized features: spectral centroid, spectral bandwidth, HNR, and pitch.

Cluster Composition

The cluster-composition panel shows how the analysis grains are distributed among clusters. Each horizontal bar gives the number of grains assigned to that cluster. Its label also reports the cluster's mean HNR in dB and mean spectral centroid in Hz. The winning cluster is explicitly marked selected; clusters with no assigned grains are shown as empty. Up to 12 requested clusters are displayed.

Grain Reuse

The synthesis draws grains randomly with replacement from the selected cluster, so some source grains may appear many times while others may never be chosen. The Grain Reuse panel makes that distribution visible:

Summary strip

The bottom strip reports four groups of run statistics:

RowWhat it reports
InputSource name and duration, grain size, effective grain crossfade, and number of analysis grains.
ClusteringK-means iteration setting, active/requested clusters, selected cluster, selected-grain count, and the number of grains with unmeasurable HNR, including how many of those entered the selected cluster and the floor value used for plotting/analysis.
SynthesisTotal rendered grains, number of distinct selected grains actually used, mean and maximum grain reuse, requested shimmer configuration, and the realised percentage/count of shimmer events.
OutputLayer count, mono/stereo rendering description, requested and rendered duration, channel count, and final measured peak.
The visualization reports the current run. Random grain selection and shimmer decisions therefore affect the Grain Reuse and realised-shimmer statistics from one run to another.

Parameters & Settings

Synthesis Parameters

ParameterTypeDefaultDescription
output_duration_secpositive20.0Duration of generated drone
layer_densitypositive3Number of parallel texture layers
add_octave_shimmerboolean1Enable harmonic pitch shifting

AI Analysis Parameters

ParameterTypeDefaultDescription
grain_size_mspositive100Duration of analysis grains
number_of_clustersinteger3K-means cluster count

Output Parameters

ParameterTypeDefaultDescription
play_resultboolean1Auto-play after generation

Parameter Interactions

Grain size effects:
  • Small (20-50ms): Very granular, abstract texture, less source recognition
  • Medium (80-150ms): Balanced, good source preservation with transformation
  • Large (200-500ms): Smooth evolution, strong source character, less granular
Layer density guidance:
  • 1 layer: Sparse, transparent, good for background
  • 2-3 layers: Rich, full, general purpose
  • 4-5 layers: Dense, complex, foreground texture
  • >5 layers: Potentially muddy, use with bright sources
Cluster count strategy:
  • 2 clusters: Broad separation (tonal vs noisy)
  • 3-4 clusters: Good detail, recommended default
  • 5-8 clusters: Fine-grained separation, for complex sources
  • >8 clusters: Over-segmentation, usually unnecessary

Applications

Ambient Music Production

Use case: Generating evolving pads and textures for ambient tracks

Technique: Use vocal or string sources with 3-4 layers, octave shimmer enabled

Workflow: Generate multiple drones from same source, layer in DAW

Soundscape Design

Use case: Creating environmental and fictional soundscapes

Technique: Use field recordings as source, small grain sizes (50ms)

Examples: Water sounds → flowing textures, forest → shimmering pads

Film and Game Audio

Use case: Background textures for cinematic scenes and game environments

Advantages:

Source Transformation

Use case: Radical transformation while preserving essence

Technique: Use distinctive sources with large grain sizes

Examples: Speech → textural clouds, percussion → rhythmic beds

Practical Workflow Examples

🎵 Ambient Pad Generation

Goal: Create evolving pad from vocal source

Settings:

  • Source: Sustained vocal note
  • Output duration: 60.0
  • Layer density: 4
  • Octave shimmer: Enabled
  • Grain size: 120ms
  • Clusters: 4

Result: Rich, evolving pad with vocal character

🎬 Cinematic Texture

Goal: Dark atmospheric texture for film

Settings:

  • Source: Low cello note
  • Output duration: 120.0
  • Layer density: 3
  • Octave shimmer: Disabled
  • Grain size: 200ms
  • Clusters: 3

Result: Dark, slowly evolving atmospheric bed

🌊 Water Transformation

Goal: Abstract water-like texture

Settings:

  • Source: Stream recording
  • Output duration: 30.0
  • Layer density: 2
  • Octave shimmer: Enabled
  • Grain size: 80ms
  • Clusters: 5

Result: Shimmering, fluid texture with water character

Advanced Techniques

Multi-stage processing:
  • Stage 1: Generate drone from source A
  • Stage 2: Use drone as source for second generation
  • Stage 3: Layer multiple generation stages
  • Result: Complex, deeply processed textures
Parameter automation:
  • Gradual density increase: Start sparse, build density
  • Grain size evolution: Large to small for increasing abstraction
  • Cluster focus shifting: Transition between different tonal groups

Create dynamic, evolving compositions rather than static textures

Troubleshooting Common Issues

Problem: Output too noisy/unmusical
Cause: Source lacks tonal content, too many clusters selecting noisy grains
Solution: Use more tonal source, reduce cluster count, increase grain size
Problem: Output too repetitive
Cause: Source too short, too few tonal grains available
Solution: Use longer source, reduce output duration, increase cluster count
Problem: Processing too slow
Cause: Very long source and/or output duration
Solution: Use shorter source excerpt, reduce output duration, increase grain size
Problem: Memory errors
Cause: Too many temporary objects, insufficient RAM
Solution: Close other Praat objects, reduce layer density, use shorter source

Algorithmic Deep Dive

Mathematical Foundations

Distance Metrics

Euclidean distance in 4D space:

Given two feature vectors: V1 = [centroid1, bandwidth1, harmonicity1, pitch1] V2 = [centroid2, bandwidth2, harmonicity2, pitch2] Distance calculation: d = √[(c1-c2)² + (b1-b2)² + (h1-h2)² + (p1-p2)²] Why Euclidean? Simple, fast computation Works well with normalized features Intuitive geometric interpretation Alternative: Manhattan distance could be used d = |c1-c2| + |b1-b2| + |h1-h2| + |p1-p2| But Euclidean generally preferred for continuous features

Convergence Criteria

K-means stopping conditions:

Primary: Assignment stability IF no grains change cluster assignment in E-step THEN algorithm converged Secondary: Iteration limit Maximum iterations = 10 (hard-coded) Prevents infinite loops Usually converges in 3-6 iterations Why 10 iterations sufficient? 4D space with normalized features Typically 100-1000 grains Convergence usually rapid Check: Count grains that change assignment each iteration When changes = 0 → converged

Computational Complexity

Time Requirements

Major processing stages:

1. FEATURE EXTRACTION: O(nGrains) Each grain: spectral analysis + HNR + pitch Most expensive: spectral analysis 2. NORMALIZATION: O(nGrains × 4) Simple arithmetic per feature per grain 3. CLUSTERING: O(iterations × nGrains × k × 4) Per iteration: for each grain, compare to k centroids Typically: 5 iterations × 500 grains × 3 clusters × 4 features 4. SYNTHESIS: O(output_duration / grain_size × layer_density) Grain concatenation and mixing Bottleneck: Feature extraction (spectral analysis) Scale factor: Source duration → more grains → longer processing

Memory Requirements

Major memory consumers:

1. ORIGINAL SOUND: duration × sampling_rate × 4 bytes Example: 60s × 44100 Hz × 4 bytes = ~10MB 2. TEMPORARY GRAINS: nGrains × grain_size × 4 bytes Example: 500 grains × 0.1s × 44100 Hz × 4 bytes = ~8MB 3. FEATURE TABLES: nGrains × 4 features × 8 bytes Example: 500 × 4 × 8 = 16KB (negligible) 4. LAYER CONSTRUCTION: layer_density × output_duration × 4 bytes Example: 3 layers × 20s × 44100 Hz × 4 bytes = ~10MB TOTAL: Approximately 2-3× original sound size Peak during synthesis: multiple layers + temporary grains