Self-Attention Recomposer — User Guide
Attention-inspired audio recomposition using MFCC chunk embeddings, query-updated similarity retrieval, optional probabilistic sampling, and configurable chunk reconstruction.
What this does
The script segments one Sound into chunks, computes one MFCC-based embedding per usable chunk, normalizes the embedding space, and then constructs a new chunk sequence. At each output step, a query vector is compared with all currently eligible chunk embeddings by dot product. Scores may be modified by temporal and repetition penalties, converted to a probability distribution, filtered, and then used for Greedy or Sampling selection.
What makes this “self-attention-inspired” rather than a Transformer?
There is a real query-to-embedding similarity mechanism, but no learned Q/K/V projections, no multi-head attention, no chunk-by-chunk attention matrix, no value aggregation, and no training. The most precise technical description is self-attention-inspired autoregressive MFCC retrieval.
The output is a recomposed sequence of source chunks. The original channel count is preserved during audio reconstruction even though MFCC analysis itself is performed on mono chunk copies.
Quick Start
- Select exactly one Sound object.
- Run
Self_Attention_Recomposer.praat. - Choose Custom or one of the four presets.
- If using Custom, set segmentation, context, selection, permutation, start mode, and join mode.
- Set Random_seed to 0 for unpredictable runs or to a positive integer for reproducible random behavior.
- Choose Fade-separated concatenation or true overlap Crossfade.
- Enable Draw_visualization and Play_result as desired.
Attention-Inspired Retrieval
Core score
Because usable chunk embeddings are L2-normalized, the dot product between two chunk embeddings is cosine similarity. Mean and EMA queries are not necessarily unit length unless Renormalize_query is enabled.
Score modifiers
Softmax
Softmax is computed only over eligible chunks.
MFCC Embeddings
MFCC extraction uses:
MFCC window: 25 ms MFCC step: 10 ms Mel bands: 24 Frequency range: 100–5000 Hz
Undefined MFCC frame values are skipped rather than inserted as zero.
Short and unanalysable chunks
The effective minimum chunk duration is at least:
If Min_chunk_duration_s is smaller, the script automatically raises the effective minimum. Chunks with no analysable MFCC frames are excluded from the retrieval pool rather than receiving an artificial zero embedding.
Normalization
- Each embedding dimension is z-scored across usable embeddings only.
- Each usable chunk embedding is then L2-normalized.
Context Modes
Last chunk
The query is automatically unit length because it is exactly one normalized chunk embedding.
Mean of last N
Exponential moving average
What does Renormalize_query change?
By default, Mean and EMA queries are not L2-renormalized. Their length therefore reflects how coherent the recent context is: aligned embeddings produce a longer query and a sharper score distribution, while disagreeing embeddings shorten the query and flatten it. Turning Renormalize_query on removes this coupling by returning the updated query to unit length.
Eligibility, Filtering, and Selection
Hard eligibility
Chunks with invalid embeddings are never eligible. In Permutation_mode, previously used chunks are also removed from eligibility completely; they receive exactly zero probability.
Top-K
When enabled, Top-K keeps only the K highest-scoring eligible chunks before softmax. In the default script it is an internal parameter; the Exploratory preset sets TopK = 5.
Top-P
When enabled, nucleus filtering retains the smallest descending-probability set whose cumulative probability reaches TopP, then renormalizes. No current preset enables TopP.
Selection modes available in the form
| Mode | Behavior |
|---|---|
| Greedy | Select the eligible chunk with the highest final weight. |
| Sampling | Draw one eligible chunk from the final categorical distribution. |
An internal centroid-pick path also exists in the script, but it is not exposed by the form and is not enabled by any current preset.
Permutation behavior
- Output_length_chunks = 0 → output length equals the number of usable chunks.
- Requested length shorter than the usable pool + Permutation ON → partial permutation.
- Requested length equal to the usable pool → full permutation.
- Requested length longer than the usable pool → permutation is disabled and repeats are allowed.
Preset Strategies
| Preset | Main settings | Important behavior |
|---|---|---|
| Smooth Flow | 0.30 s chunks; EMA α=0.8; Greedy; near-dup 0.3; time decay τ=2.0; permutation ON; T=0.5 | Temperature does not change its Greedy ordering. |
| Jumpy Mosaic | 0.15 s chunks; Last chunk; Sampling; near-dup 0.8; permutation ON; T=1.5 | Temperature affects the sampling distribution. |
| Exploratory | 0.20 s chunks; Mean of last 6; Sampling; variance ON; TopK=5; repeats allowed; repeat penalty 1.5; T=1.2 | Uses 26-dimensional embeddings. |
| Strict Permutation | 0.25 s chunks; Last chunk; Greedy; near-dup 0.5; permutation ON; T=0.3 | Uses each usable chunk at most once; Temperature does not change Greedy ordering. |
Audio Reconstruction
Selected chunks are re-extracted from the original Sound, so the source channel count is preserved in the rendered output.
Fade-separated concatenation vs. real overlap crossfade
Fade-separated concatenation — default: every chunk receives a linear fade-in and fade-out, then the chunks are joined with ordinary Concatenate. There is no temporal overlap between adjacent chunks.
Crossfade: per-chunk fades are skipped, and Praat's Concatenate with overlap is used instead. The requested overlap is Fade_duration_s, capped at 45% of the shortest selected chunk.
If the output contains only one chunk, no join operation is performed and diagnostics report:
single chunk (no join)
Output naming
[OriginalName]_attnRec_[PresetName]
Parameters
| Parameter | Default | Behavior |
|---|---|---|
| Preset | Custom | Custom, Smooth Flow, Jumpy Mosaic, Exploratory, Strict Permutation. |
| Use_silence_segmentation | Off | Use silence-based TextGrid segmentation instead of fixed-length chunks. |
| Chunk_duration_s | 0.20 | Fixed segmentation duration when silence segmentation is off. |
| Min_chunk_duration_s | 0.05 | Minimum chunk duration; internally raised to at least the MFCC analysis minimum. |
| Temperature | 1.0 | Clamped to at least 0.01. Affects Sampling probabilities; not Greedy ordering. |
| Context_mode | Last chunk | Last chunk, Mean of last N, or Exponential moving average. |
| Renormalize_query | Off | L2-renormalize Mean/EMA queries after every update. |
| Selection_mode | Greedy | Greedy or Sampling. |
| Permutation_mode | On | Prevent reuse of eligible chunks while enough unused chunks remain. |
| Output_length_chunks | 0 | 0 = number of usable chunks; positive values request a specific output length. |
| Start_mode | First chunk | First usable, Random usable, or Highest-RMS usable chunk. |
| Chunk_join | Fade-separated concatenation | Fade-separated butt join or true overlap crossfade. |
| Fade_duration_s | 0.01 | Per-chunk fade duration or requested overlap duration, depending on Chunk_join. |
| Random_seed | 0 | 0 = unpredictable; positive integer = reproducible Random start/Sampling behavior. |
| Draw_visualization | On | Draw the suite-standard analysis display. |
| Play_result | On | Play the final output after processing. |
Internal parameters
| Parameter | Default |
|---|---|
| num_coefficients | 13 |
| mfcc_window_s | 0.025 |
| mfcc_step_s | 0.010 |
| max_frequency_Hz | 5000 |
| use_variance | 0 |
| context_window | 4 |
| ema_alpha | 0.7 |
| top_k | 0 |
| top_p | 0 |
| repeat_penalty | 2.0 |
| near_dup_penalty | 0.5 |
| use_time_decay | 0 |
| time_tau | 1.0 |
| silence_threshold_dB | -25 |
| min_silence_duration_s | 0.10 |
Visualization
The current v1.3 visualization uses the Praat AudioTools suite layout:
- Header — input, preset, chunk/step count, Temperature, context and selection mode.
- Original waveform — source waveform with chunk boundaries.
- Output waveform — reconstructed Sound.
- Attention order path — source chunk index versus output step.
- Consecutive similarity — dot-product similarity between consecutive selected chunk embeddings, with a mean reference line.
- Summary strip — embedding dimension, permutation/repeat mode, context, selection, Temperature, TopK/TopP, entropy, average similarity, join description, fade/overlap duration, and sample rate.
There is no standalone legend panel or statistics panel in v1.3; that information is consolidated into the header and summary strip.
Interpretation and Scope
The script is suited to deterministic or stochastic chunk recomposition, timbre-oriented retrieval, remix-like permutations, and generative texture construction.