Self-Attention Recomposer — User Guide

Attention-inspired audio recomposition using MFCC chunk embeddings, query-updated similarity retrieval, optional probabilistic sampling, and configurable chunk reconstruction.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 1.3 (2026) Technique: Self-attention-inspired autoregressive MFCC retrieval Category: Composition
Contents:

What this does

The script segments one Sound into chunks, computes one MFCC-based embedding per usable chunk, normalizes the embedding space, and then constructs a new chunk sequence. At each output step, a query vector is compared with all currently eligible chunk embeddings by dot product. Scores may be modified by temporal and repetition penalties, converted to a probability distribution, filtered, and then used for Greedy or Sampling selection.

What makes this “self-attention-inspired” rather than a Transformer?

There is a real query-to-embedding similarity mechanism, but no learned Q/K/V projections, no multi-head attention, no chunk-by-chunk attention matrix, no value aggregation, and no training. The most precise technical description is self-attention-inspired autoregressive MFCC retrieval.

The output is a recomposed sequence of source chunks. The original channel count is preserved during audio reconstruction even though MFCC analysis itself is performed on mono chunk copies.

Quick Start

  1. Select exactly one Sound object.
  2. Run Self_Attention_Recomposer.praat.
  3. Choose Custom or one of the four presets.
  4. If using Custom, set segmentation, context, selection, permutation, start mode, and join mode.
  5. Set Random_seed to 0 for unpredictable runs or to a positive integer for reproducible random behavior.
  6. Choose Fade-separated concatenation or true overlap Crossfade.
  7. Enable Draw_visualization and Play_result as desired.
Presets override the corresponding form values after the dialog is closed. To make fully manual choices, use Custom.

Attention-Inspired Retrieval

Core score

score_i = query · embedding_i

Because usable chunk embeddings are L2-normalized, the dot product between two chunk embeddings is cosine similarity. Mean and EMA queries are not necessarily unit length unless Renormalize_query is enabled.

Score modifiers

optional time decay: score_i -= |chunkMid_i - chunkMid_current| / time_tau repeat penalty: score_i -= repeat_penalty × useCount_i (only when repeats are allowed) near-duplicate penalty: score_i -= near_dup_penalty × similarity(currentChunk, candidate)
Important: the near-duplicate penalty compares each candidate with the last selected chunk embedding, not with the current query. This matters when the query represents a mean or EMA of several previous chunks.

Softmax

weight_i = exp((score_i - maxScore) / Temperature) / sum(exp(...))

Softmax is computed only over eligible chunks.

Temperature does not change Greedy ordering. For any positive Temperature, the largest softmax weight is attached to the same chunk as the largest score. In Greedy mode, Temperature changes only the displayed weights and entropy. In Sampling mode, it changes the probability distribution and therefore can change the audio result.

MFCC Embeddings

What represents each chunk?

Each usable chunk is converted to mono for analysis and represented by the mean of 13 MFCC coefficients. When variance is enabled, 13 additional within-chunk variance values are appended, producing a 26-dimensional embedding.

MFCC extraction uses:

MFCC window: 25 ms
MFCC step:   10 ms
Mel bands:   24
Frequency range: 100–5000 Hz

Undefined MFCC frame values are skipped rather than inserted as zero.

Short and unanalysable chunks

The effective minimum chunk duration is at least:

1.5 × MFCC window = 1.5 × 25 ms = 37.5 ms

If Min_chunk_duration_s is smaller, the script automatically raises the effective minimum. Chunks with no analysable MFCC frames are excluded from the retrieval pool rather than receiving an artificial zero embedding.

Normalization

  1. Each embedding dimension is z-scored across usable embeddings only.
  2. Each usable chunk embedding is then L2-normalized.
z_d = (x_d - mean_d) / std_d e = z / ||z||₂

Context Modes

Last chunk

query_(t+1) = embedding(chosen_t)

The query is automatically unit length because it is exactly one normalized chunk embedding.

Mean of last N

query_(t+1) = mean of the most recent context_window embeddings

Exponential moving average

query_(t+1) = ema_alpha × embedding(chosen_t) + (1 - ema_alpha) × query_t

What does Renormalize_query change?

By default, Mean and EMA queries are not L2-renormalized. Their length therefore reflects how coherent the recent context is: aligned embeddings produce a longer query and a sharper score distribution, while disagreeing embeddings shorten the query and flatten it. Turning Renormalize_query on removes this coupling by returning the updated query to unit length.

Eligibility, Filtering, and Selection

Hard eligibility

Chunks with invalid embeddings are never eligible. In Permutation_mode, previously used chunks are also removed from eligibility completely; they receive exactly zero probability.

Top-K

When enabled, Top-K keeps only the K highest-scoring eligible chunks before softmax. In the default script it is an internal parameter; the Exploratory preset sets TopK = 5.

Top-P

When enabled, nucleus filtering retains the smallest descending-probability set whose cumulative probability reaches TopP, then renormalizes. No current preset enables TopP.

Selection modes available in the form

ModeBehavior
GreedySelect the eligible chunk with the highest final weight.
SamplingDraw one eligible chunk from the final categorical distribution.

An internal centroid-pick path also exists in the script, but it is not exposed by the form and is not enabled by any current preset.

Permutation behavior

Preset Strategies

PresetMain settingsImportant behavior
Smooth Flow 0.30 s chunks; EMA α=0.8; Greedy; near-dup 0.3; time decay τ=2.0; permutation ON; T=0.5 Temperature does not change its Greedy ordering.
Jumpy Mosaic 0.15 s chunks; Last chunk; Sampling; near-dup 0.8; permutation ON; T=1.5 Temperature affects the sampling distribution.
Exploratory 0.20 s chunks; Mean of last 6; Sampling; variance ON; TopK=5; repeats allowed; repeat penalty 1.5; T=1.2 Uses 26-dimensional embeddings.
Strict Permutation 0.25 s chunks; Last chunk; Greedy; near-dup 0.5; permutation ON; T=0.3 Uses each usable chunk at most once; Temperature does not change Greedy ordering.

Audio Reconstruction

Selected chunks are re-extracted from the original Sound, so the source channel count is preserved in the rendered output.

Fade-separated concatenation vs. real overlap crossfade

Fade-separated concatenation — default: every chunk receives a linear fade-in and fade-out, then the chunks are joined with ordinary Concatenate. There is no temporal overlap between adjacent chunks.

Crossfade: per-chunk fades are skipped, and Praat's Concatenate with overlap is used instead. The requested overlap is Fade_duration_s, capped at 45% of the shortest selected chunk.

If the output contains only one chunk, no join operation is performed and diagnostics report:

single chunk (no join)

Output naming

[OriginalName]_attnRec_[PresetName]

Parameters

ParameterDefaultBehavior
PresetCustomCustom, Smooth Flow, Jumpy Mosaic, Exploratory, Strict Permutation.
Use_silence_segmentationOffUse silence-based TextGrid segmentation instead of fixed-length chunks.
Chunk_duration_s0.20Fixed segmentation duration when silence segmentation is off.
Min_chunk_duration_s0.05Minimum chunk duration; internally raised to at least the MFCC analysis minimum.
Temperature1.0Clamped to at least 0.01. Affects Sampling probabilities; not Greedy ordering.
Context_modeLast chunkLast chunk, Mean of last N, or Exponential moving average.
Renormalize_queryOffL2-renormalize Mean/EMA queries after every update.
Selection_modeGreedyGreedy or Sampling.
Permutation_modeOnPrevent reuse of eligible chunks while enough unused chunks remain.
Output_length_chunks00 = number of usable chunks; positive values request a specific output length.
Start_modeFirst chunkFirst usable, Random usable, or Highest-RMS usable chunk.
Chunk_joinFade-separated concatenationFade-separated butt join or true overlap crossfade.
Fade_duration_s0.01Per-chunk fade duration or requested overlap duration, depending on Chunk_join.
Random_seed00 = unpredictable; positive integer = reproducible Random start/Sampling behavior.
Draw_visualizationOnDraw the suite-standard analysis display.
Play_resultOnPlay the final output after processing.

Internal parameters

ParameterDefault
num_coefficients13
mfcc_window_s0.025
mfcc_step_s0.010
max_frequency_Hz5000
use_variance0
context_window4
ema_alpha0.7
top_k0
top_p0
repeat_penalty2.0
near_dup_penalty0.5
use_time_decay0
time_tau1.0
silence_threshold_dB-25
min_silence_duration_s0.10

Visualization

The current v1.3 visualization uses the Praat AudioTools suite layout:

  1. Header — input, preset, chunk/step count, Temperature, context and selection mode.
  2. Original waveform — source waveform with chunk boundaries.
  3. Output waveform — reconstructed Sound.
  4. Attention order path — source chunk index versus output step.
  5. Consecutive similarity — dot-product similarity between consecutive selected chunk embeddings, with a mean reference line.
  6. Summary strip — embedding dimension, permutation/repeat mode, context, selection, Temperature, TopK/TopP, entropy, average similarity, join description, fade/overlap duration, and sample rate.
How to read the order path: the vertical coordinate is the original source-chunk index and the horizontal coordinate is the generated output step. Large vertical jumps indicate movement to distant source chunks, not necessarily large timbral changes.
How to read consecutive similarity: because the chunk embeddings themselves are L2-normalized, this diagnostic is the dot product between consecutive chunk embeddings. It is therefore a cosine-similarity measure in the normalized MFCC space.

There is no standalone legend panel or statistics panel in v1.3; that information is consolidated into the header and summary strip.

Interpretation and Scope

The script is suited to deterministic or stochastic chunk recomposition, timbre-oriented retrieval, remix-like permutations, and generative texture construction.

The mechanism does not learn musical syntax or statistical patterns from a dataset. It operates only on the chunks of the currently selected Sound and derives each next-step distribution directly from the current query, the source embeddings, and the configured penalties and filters.