Granular Attention Re-synthesis — User Guide

Attention-weighted self-resynthesis: candidate grains are scored, hard-gated, and sampled from an active-only softmax distribution. Low alpha gives a broad self-remix; high alpha concentrates repeatedly on the strongest surviving grains.

Author: Shai Cohen Version: 2.1 (2026) Technique: Attention-Based Grain Selection Category: Granular / Composition / Experimental Citation: Cohen, S. (2026). Praat AudioTools
Contents:

What this does

This script implements a Granular Attention Re-synthesis engine in which candidate grains compete for being selected, not for gain. The source is converted to mono, analyzed as a bank of overlapping candidate grains, and reconstructed from those same moments according to an attention distribution derived from energy and/or energy-change features.

🎯 What is attention-based grain selection?

  • Candidate grains are extracted from the mono source at regular candidate-hop intervals.
  • Each candidate is Hann-windowed for analysis, matching the windowed content that reaches the resynthesis path.
  • Each candidate receives a score: RMS energy, onset slope, energy-edge slope, or a mixed score.
  • Hard attention gate: candidates below the floor are removed from competition entirely.
  • Active-only softmax: only surviving grains receive non-zero probability.
  • Attention_sharpness_alpha: higher values sharpen the probability distribution; lower values flatten it.
  • Recency penalty: the last three selected grains are attenuated before each draw and the distribution is renormalized.
  • True normalized Hann OLA: grains are placed in a buffer at real output times, with an accumulated Hann envelope used for normalization.

The result is a self-remix that can move from broad textural redistribution to concentrated self-quotation and stutter.

Key Features:

Technical implementation: (1) convert source to mono and keep an unmodified-level dry reference; (2) peak-normalize a working mono copy; (3) extract Hann-windowed candidate grains and score them; (4) apply a hard active mask and softmax over survivors only; (5) repeatedly draw grains with recency-adjusted probabilities; (6) apply optional varispeed before the Hann window; (7) place each grain at its actual jittered onset into output and envelope buffers; (8) divide by accumulated envelope; (9) trim/fade to source duration; (10) blend wet with the dry source; (11) peak-scale the final mono output to 0.99.

Quick start

  1. In Praat, select exactly one Sound object, at least 50 ms long.
  2. Run script… → Granular_Attention_Resynth.praat.
  3. Choose a preset. Self-Remix is a good starting point.
  4. For Custom, set grain size, synthesis hop, candidate hop, attention sharpness, floor, score type, jitter, recency, wet percentage, and Random_seed.
  5. Enable visualization if you want to inspect the attention landscape and actual output-to-source read map.
  6. Click OK. The result appears as <source>_GAR_<preset>.
Quick tip: Start with Self-Remix. Lower alpha broadens the active distribution; higher alpha makes a few high-scoring grains dominate. Raise Recency_penalty when concentrated presets repeat the same grains too often. Use Random_seed > 0 when you want to reproduce a successful take.
Important: The processor is mono by design. Stereo or multichannel input is summed to mono. Grain size is clamped to 5–1000 ms and never allowed to exceed source duration. Synthesis hop is limited to at most half the effective grain duration; candidate hop cannot exceed the grain duration; time jitter cannot exceed one synthesis hop; pitch jitter is clamped to 0–4 semitones.

Attention Selection Theory

The Attention Gate and Softmax

🧮 Active-only attention distribution

meanScore = (1/N) * sum(rawScore[i]) floorFactor = 10^(Floor_dB / 10) floorValue = meanScore * floorFactor ACTIVE MASK: if rawScore[i] >= floorValue: active[i] = 1 gated[i] = rawScore[i] else: active[i] = 0 probability[i] = 0 If no grain survives: activate only the single highest-scoring grain ACTIVE-ONLY SOFTMAX: maxGated = max(gated among active grains) weight[i] = exp(((gated[i] - maxGated) / meanScore) * alpha) for active grains only probability[i] = weight[i] / sum(active weights) Rejected grains remain exactly 0.

Important terminology: the control is Attention_sharpness_alpha, not “temperature”. Higher alpha sharpens the distribution. In conventional softmax notation softmax(z/T), higher temperature would do the opposite.

Score Types

📊 Four Scoring Methods

Score TypeDefinitionMusical tendency
RMSHann-windowed RMS², normalized by maximum RMS²Louder candidate regions compete more strongly
OnsetsPositive change in RMS² only, normalized by maximum positive slopeRising-energy events and attacks are favored
Energy edgesAbsolute change in RMS², normalized by maximum absolute slopeBoth attacks and decays can be emphasized
Mixed(1-w) * RMS_norm + w * positiveSlope_normBlend of energy and onset emphasis

Mixed mode uses the positive-slope transient measure, not the absolute energy-edge score.

Recency Penalty

Before EACH random draw: if candidate == lastGrain1: p *= (1 - recency_penalty) if candidate == lastGrain2: p *= (1 - recency_penalty * 0.6) if candidate == lastGrain3: p *= (1 - recency_penalty * 0.3) Then renormalize the penalized distribution and draw one grain by inverse-transform sampling.

The blue probability curve in the visualization shows the base active-only softmax. Actual usage also reflects the per-draw recency penalty, so usage does not have to match the blue line exactly.

Musical Effects by Attention Sharpness α

α RangeBehaviorMusical ResultExample
1–3Broad competitionTextural self-remix; many active grains remain plausibleSelf-Remix, Shimmer
5–10Moderate concentrationHigh-score moments dominate more stronglyOnset Harvest, Slabs
12+Strong concentrationCrystallization, self-quotation, repeated motifsCrystallize, Motif Extract

True Hann Overlap-Add

For each output hop: 1. Draw a candidate grain. 2. Extract it RECTANGULARLY from the mono source. 3. Optional varispeed: shiftFactor = 2^(pitchShift / 12) resample, then override sampling frequency to source SR 4. Apply ONE Hann window. 5. Nominal output position = (hop - 1) * synthHop 6. Add random output-position jitter in +/- Time_jitter 7. Accumulate: outBuf += windowed grain envBuf += the same Hann window After all grains: envFloor = 0.15 * peak(envBuf) wet = outBuf / max(envBuf, envFloor) Then: trim to source duration apply short head/tail fades wet/dry blend final peak scale to 0.99
Why this matters: v2.0 removed the old “Hann grain + Concatenate with overlap” architecture. That older approach effectively applied a second crossfade envelope and produced repeated level dips at joins. The current buffer OLA supports more than two simultaneous grains and normalizes by the accumulated Hann envelope.

Varispeed Pitch Jitter

Pitch_jitter_semitones is varispeed. It is not a duration-preserving pitch shifter. Pitch, playback speed, and grain duration move together. The varispeed operation happens before the Hann window, and grains are not forced back to one common duration.

Preset Strategies

Preset 2: Self-Remix

🌱 Gentle Textural Remix

Grain: 60 ms | Hop: 30 ms | Candidate: 20 ms

Alpha: 3.0 | Floor: -3 dB | Score: RMS

Time jitter: ±8 ms | Pitch: 0 st | Recency: 0.3 | Wet: 100%

Preset 3: Crystallize

💎 Crystallized Texture

Grain: 40 ms | Hop: 20 ms | Candidate: 15 ms

Alpha: 12.0 | Floor: +3 dB | Score: RMS

Time jitter: ±3 ms | Pitch: 0 st | Recency: 0.4 | Wet: 100%

Preset 4: Motif Extract

🔁 Strong Motif Concentration

Grain: 30 ms | Hop: 15 ms | Candidate: 10 ms

Alpha: 25.0 | Floor: +6 dB | Score: RMS

Time jitter: ±1 ms | Pitch: 0 st | Recency: 0.6 | Wet: 100%

Preset 5: Onset Harvest

⚡ Positive-Slope Attack Focus

Grain: 50 ms | Hop: 25 ms | Candidate: 15 ms

Alpha: 8.0 | Floor: +2 dB | Score: Onsets (positive slope only)

Time jitter: ±5 ms | Pitch: 0 st | Recency: 0.4 | Wet: 100%

Preset 6: Shimmer

✨ Shimmering Texture

Grain: 20 ms | Hop: 10 ms | Candidate: 10 ms

Alpha: 2.0 | Floor: -6 dB | Score: RMS

Time jitter: ±4 ms | Varispeed: ±0.3 st | Recency: 0.2 | Wet: 85%

Preset 7: Slabs

🧱 Large-Grain Mosaic

Grain: 300 ms | Hop: 150 ms | Candidate: 50 ms

Alpha: 10.0 | Floor: +2 dB | Score: RMS

Time jitter: ±15 ms | Pitch: 0 st | Recency: 0.5 | Wet: 100%

Preset 8: Cloud

☁️ Drifting Cloud

Grain: 600 ms | Hop: 300 ms | Candidate: 80 ms

Alpha: 4.0 | Floor: -3 dB | Score: Mixed (transient weight 0.4)

Time jitter: ±30 ms | Pitch: 0 st | Recency: 0.3 | Wet: 90%

Preset scope: Built-in presets override the grain/hop, alpha, floor, score, transient weight, jitter, recency, and wet controls listed above. Random_seed, Draw_visualization, and Play_result are not overridden.

Parameters & Controls

Custom Defaults

ParameterDefaultValidation / meaning
Grain_size_ms150.0Clamped to 5–1000 ms and then to source duration
Synthesis_hop_ms50.0Minimum 1 ms; capped to half the effective grain duration; may be raised to keep total hops ≤ 20,000
Candidate_hop_ms20.0Minimum 1 ms; capped to effective grain duration
Attention_sharpness_alpha1.0Internal minimum 0.01; higher = sharper active distribution
Floor_dB0.0Threshold relative to mean score in dB: mean × 10^(dB/10)
Score_typeRMSRMS / Onsets / Energy edges / Mixed
Transient_weight0.5Clamped to 0–1; used only by Mixed
Time_jitter_ms5.0Negative → 0; capped to one effective synthesis hop
Pitch_jitter_semitones0.0Clamped to 0–4 st; varispeed, not duration-preserving pitch shift
Recency_penalty0.5Clamped to 0–1
Wet_percent100.0Clamped to 0–100
Random_seed00 = unpredictable; positive integer = reproducible random draws
Draw_visualization1Draw v2.1 diagnostic page
Play_result1Audition final output

Wet / Dry Behavior

working mono source: peak-normalized to 0.99 for candidate analysis/resynthesis dry reference: copied BEFORE that peak normalization therefore Wet_percent = 0 preserves the mono source at its original level before final output normalization mix: output = wet * Wet_percent/100 + dry * (1 - Wet_percent/100) final: entire result peak-scaled to 0.99
Important: Wet_percent = 0 does not produce a bit-identical multichannel bypass because the processor is mono and the final result is peak-scaled. It does, however, use the unmodified-level mono source as the dry signal rather than the internally peak-normalized working copy.

Visualization & Analysis

v2.1 Diagnostic Page

HEADER script version, source, preset, score type, alpha, grains used SOURCE WAVEFORM mono analysis/resynthesis path OUTPUT WAVEFORM mono final result on the same amplitude scale ATTENTION LANDSCAPE dark score stems = active candidates light grey stems = rejected candidates amber dashed line = ReLU-inspired hard gate threshold blue curve = base active-only softmax probability orange circles = actual usage; size = selection count ATTENTION RESYNTHESIS MAP x-axis = actual jittered output onset y-axis = selected source-grain start time dashed diagonal = identity/proportional read reference marker colour follows attention score plot is thinned to about 600 events when necessary USAGE DISTRIBUTION one bar per candidate grain bar colour follows score dashed line = uniform-use mean SUMMARY STRIP source duration/channels -> mono grain size / candidate hop / candidate count score, alpha, floor, active count recency penalty synthesis hop, real time jitter, varispeed range wet percentage, percentage of grains used normalized Hann OLA
Reading the visualization: A candidate can have a high raw score but zero probability if it falls below the active gate. The blue line is the unpenalized base distribution; orange usage reflects the stochastic draws after recency attenuation. The resynthesis map is the most direct view of the transformation: it shows exactly where each grain was placed in the output and where that grain came from in the source.

Applications

Generative Self-Remix

Use Self-Remix or Cloud to redistribute the source's own salient regions across its original duration. The script does not extend duration: the wet path is trimmed to source length.

Rhythmic & Glitch Effects

Shimmer and Varispeed Texture

Shimmer adds a small varispeed range before the grain window is applied. This changes pitch and temporal scale together and can create a lightweight drifting texture without pretending to be a phase-vocoder pitch shifter.

Source Navigation / Self-Quotation

The v2.1 resynthesis map makes the process useful analytically as well as compositionally: diagonal clusters indicate near-identity reading, while horizontal or repeated source-time bands reveal recurrent self-quotation.

Research & Education

The visualization can demonstrate active gating, softmax concentration, recency-modified sampling, stochastic usage, and overlap-add normalization. For conceptual accuracy, describe this as an attention-inspired grain-selection mechanism, not as a neural attention model.

Troubleshooting Common Issues

Only one grain is ever active:
Floor_dB may be too high. If no candidate passes the threshold, the script intentionally activates only the single strongest grain.
Output is too repetitive:
Reduce Attention_sharpness_alpha, lower Floor_dB, and/or increase Recency_penalty.
Output is too diffuse:
Raise alpha or raise the floor so fewer candidates remain active.
Processing is slow:
Increase Candidate_hop_ms to reduce the candidate pool and/or increase Synthesis_hop_ms to reduce output events. Decreasing either hop creates more work, not less.
Unexpected pitch-jitter timing:
Pitch jitter is varispeed. Downward shifts lengthen grains; upward shifts shorten them. The buffer OLA accepts variable-length grains directly.
Stereo source became mono:
This is intentional in the current script. Stereo and multichannel input are summed to mono before analysis/resynthesis, and the output is mono.

Advanced Techniques

Custom score functions:

The current score block can be adapted to other features, but the shipped v2.1 implementation uses only energy, positive energy slope, absolute energy slope, or their defined mixture.

Time-varying alpha:

If you modify the script, remember the correct direction: low alpha = broad/exploratory; high alpha = sharp/crystallized.