Scattering Texture Generator — Wavelet-Scattering Texture Resynthesis

Generate a new waveform that preserves selected multiscale properties of a source — coarse energy, wavelet-band spectral structure, and second-order temporal modulation — while releasing exact waveform, phase, pitch identity, and fine timing. The process combines Kymatio Morlet scattering, structural target transformation, calibrated direct synthesis, and optional iterative refinement.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 0.3.1 (2026) Category: Hybrid Systems License: MIT License Repo: Praat AudioTools
Contents:

What this does

Scattering Texture Generator analyses a selected Sound with a one-dimensional Morlet wavelet-scattering representation and constructs a new waveform realization whose selected scattering statistics approach a source-derived or transformed target. Instead of copying the original waveform, the system preserves higher-level relationships such as spectral-band energy, coarse envelope shape, and multiscale modulation behavior.

What makes this different? This is not an inverse wavelet transform, a spectral morph, a vocoder, or a neural generator. Wavelet scattering is used as a set of structural constraints. A new waveform is built from band carriers and modulation envelopes, calibrated against the target, and — in Standard and High modes — optionally refined by L-BFGS texture matching.

Key Features:

Place in the Praat AudioTools wavelet line:
CWT Scalogram — what multiscale components exist?
CWT Granular Resampler — use that structure to control another sound.
Scattering Texture Generator — generate a new waveform that preserves selected multiscale relationships.

Quick start

  1. Select exactly one Sound object in Praat.
  2. Run Scattering_Texture_Generator.praat.
  3. For a first run, leave Preset = Custom, Texture scale = Medium, and Frequency detail = Medium.
  4. Use the default preservation weights: Global envelope 40%, Spectral structure 60%, Temporal modulation 100%.
  5. Use Texture transformation to determine how strongly the scattering target itself is altered.
  6. Use Spectral randomization to progressively remove fine spectral / pitch identity from the carriers.
  7. Start with Preview for a fast direct-synthesis result. Use Standard or High when you want iterative refinement.
  8. Keep Draw visualization enabled to inspect source, target, achieved scattering structure, and loss behavior.
Good source material: sustained instrumental sounds, vocal textures, tremolo, noisy or multiphonic material, repeated attacks, granular textures, bowed sounds, percussion tails, and sounds whose identity is carried by both spectrum and temporal modulation. Sources with clear internal movement make the difference between order 1 and order 2 especially easy to hear.

How it works

1. Praat exports a mono analysis signal

The selected Sound is preserved in the Object list. If it has more than one channel, Praat converts a temporary copy to mono for analysis and synthesis, applies a protective pre-scale only when required for 24-bit WAV export, and sends the temporary WAV to the Python backend.

2. Kymatio Morlet filters define the scattering representation

The Python engine uses Kymatio's one-dimensional Morlet filter bank. The forward scattering path is implemented in NumPy with Fourier-domain filtering, dyadic subsampling, modulus, and low-pass averaging. Orders 1 and 2 follow the Kymatio scattering structure; order 0 uses an envelope form appropriate for audio.

Order 0: S0 = |x| * phi_J Order 1: S1 = |x * psi_1| * phi_J Order 2: S2 = ||x * psi_1| * psi_2| * phi_J Normalized order 2: S2_normalized = S2 / S1(parent band)

3. The source representation becomes a texture target

The source coefficients are moved into a log-domain target representation. A preset may apply its own structural operations, while Texture transformation scales a second group of target operations. These transformations act on the representation itself rather than functioning as a dry/wet control.

4. A new waveform is built from the target

For every first-order band, the engine creates a phase-randomized carrier. Its magnitude spectrum ranges from the source's detailed spectrum toward a smoothed pitch-reduced envelope according to Spectral randomization. Time-varying first-order envelopes and second-order Morlet-band modulation signals are then imposed on those carriers.

5. Calibration and optional refinement

The direct-synthesis stage repeatedly re-analyses the new waveform and corrects first-order levels and second-order modulation depths. Standard and High modes then use an analytic scattering gradient and L-BFGS to reduce a weighted log-domain mismatch.

Loss = w0 · d(S0) + w1 · d(S1) + w2 · d(S2 / S1) where d is a mean-squared log-domain distance computed separately for each coefficient family.
Not inverse scattering. The scattering transform is not directly inverted. The output is a new waveform that satisfies the selected texture constraints approximately. This is why two different random seeds can produce different sounds that belong to the same target texture family.

Scattering orders — Musical interpretation

Order 0 — Global envelope

Representation: |x| * phi_J

Musical role: coarse amplitude shape and slow energy evolution.

Control: Preserve global envelope.

Order 1 — Spectral structure

Representation: averaged magnitudes of first-order Morlet bands.

Musical role: broad spectral distribution, band energy, and timbral trajectory.

Control: Preserve spectral structure.

Order 2 — Temporal modulation

Representation: modulation of the first-order band envelopes, matched in normalized S2/S1 form.

Musical role: pulsation, tremolo-like behavior, articulation density, roughness, and multiscale texture motion.

Control: Preserve temporal modulation.

What scattering adds beyond the CWT tools: a normal CWT scalogram shows first-order time-frequency energy. Second-order scattering explicitly represents modulation inside those wavelet-band envelopes as a separate layer that can be preserved, transformed, and matched during synthesis.

6 Presets + Custom

Custom

Weights: user controls.

Transformation family: order-1 time smoothing, modulation-depth emphasis, and order-2 sparsification, scaled by Texture transformation.

Use: direct manual control of the three scattering-order weights.

Spectral Skeleton

Weights 0/1/2: 30 / 100 / 15.

Transformation: order-1 time smoothing and order-2 sparsification.

Use: retain broad timbral structure while loosening temporal modulation.

Temporal Texture

Weights 0/1/2: 30 / 35 / 100.

Preset operation: mild order-1 frequency blur. Transformation adds further blur and a downward spectral-envelope shift.

Use: foreground modulation behavior while reducing literal timbral identity.

Modulation Ghost

Weights 0/1/2: 60 / 10 / 100.

Preset operations: strong spectral blur and order-1 dynamic compression. Transformation increases modulation depth.

Use: create a temporal / textural shadow of the source.

Structure Without Identity

Weights 0/1/2: 50 / 50 / 50.

Preset operations: spectral blur and time smoothing. Transformation adds structured perturbation and an upward spectral-envelope shift.

Use: preserve broad multiscale behavior while weakening recognisable source identity.

Second-Order Reconstruction

Weights 0/1/2: 20 / 25 / 100.

Transformation: increased modulation depth.

Use: make second-order temporal structure the dominant reconstruction constraint.

Radical Texture

Weights 0/1/2: 40 / 60 / 100.

Transformation: modulation-rate shift, spectral-envelope shift, order-1 dynamic expansion, order-2 sparsification, and structured perturbation.

Use: produce a clearly transformed texture that still derives from the source's multiscale organization.

Preset behavior: named presets replace the three Preserve weights. Some presets also contain preset operations that always apply, even when Texture transformation is 0%. In the Info window and figure these are labelled [preset, always]; operations scaled by the Transformation control are labelled [x Transformation].

Controls

ControlDefaultFunction
PresetCustomSelects Custom or one of six named texture strategies. Named presets set their own order-0 / order-1 / order-2 preservation weights.
Texture scaleMedium (~186 ms)Sets the scattering averaging scale. Fine is approximately 46 ms; Medium 186 ms; Broad 743 ms. The exact J is derived from the working sample rate.
Frequency detailMedium (Q = 8)Controls first-order frequency resolution. Low uses Q=4, Medium Q=8, High Q=12; High also uses a denser second-order setting.
Preserve global envelope40%Weights order 0, preserving coarse temporal energy behavior.
Preserve spectral structure60%Weights order 1, preserving wavelet-band energy and timbral structure.
Preserve temporal modulation100%Weights normalized order 2, preserving modulation depth and multiscale textural movement.
Texture transformation30%Scales the preset's transformable operations. This changes the scattering target itself; it is not a dry/wet mix.
Spectral randomization50%Blends carrier magnitude from the source's detailed spectrum toward a smoothed pitch-reduced spectral envelope. Carrier phase is randomized at every setting.
Reconstruction qualityPreviewSelects internal working-rate cap, segment size, and optional L-BFGS refinement.
Random seed1Positive values reproduce the same realization. 0 requests a new seed for each run.
Draw visualizationOnCreates the explanatory Praat Picture display after synthesis.
PlayOnPlays the generated Sound after it is imported and named.

Reconstruction quality

ModeInternal rate capCalibrationL-BFGS refinementCharacter
Preview16 kHz5 roundsNoneFastest. Uses calibrated direct synthesis only.
Standard22.05 kHz5 roundsUp to 15 iterations per segmentBalances speed and scattering-domain refinement.
High48 kHz5 roundsUp to 150 iterations per segmentHighest working bandwidth and deepest iterative matching.
Working rate vs. output rate. Processing may occur at a lower internal rate according to the selected quality. The result is resampled back to the source sample rate and exact source sample count before Praat imports it. The Info window reports both the working rate and resulting processing band limit.

Visualisation

When Draw visualization is enabled, the script produces a Praat Picture display organised as a five-stage explanation of the transformation:

1 — Source

Shows the mono analysis copy of the selected Sound and establishes the input waveform used by the backend.

2 — Multiscale analysis

Compares source, target, and achieved low-order scattering information. Grey represents the source, blue the transformed target, and red the new realization.

3 — What scattering adds

Displays order-2 modulation across carrier band and modulation rate, making explicit the layer that is not represented separately in a standard CWT scalogram.

4 — Texture constraints + iterative matching

Shows the active order weights and the evolution of the texture-matching loss, connecting the user controls to the reconstruction process.

5 — New realization

Shows the resulting waveform together with a summary of scattering geometry, target operations, sample-rate behavior, segmentation, refinement, seed, and warnings.

The visualisation explains the transformation. It is designed to show the difference between first-order multiscale structure and second-order modulation, and to make clear that the final Sound is a newly synthesized realization rather than an inverse transform of the source.

Technical behavior

Requirements & installation

Python dependencies required: numpy, scipy, and kymatio.

Install with:
python -m pip install numpy scipy kymatio
ComponentRequirement
PraatPraat 6.1+; the script notes testing with 6.1.38, 6.4.06, and 7.0. Praat 7 may ask for full trust because the tool writes temporary files and launches Python.
PythonPython 3. The frontend uses the library's OS-specific Python discovery convention.
Python backendPlace scattering_texture_engine.py in plugin_AudioTools/py/. The script also accepts the engine next to the Praat script as a fallback.
KymatioThe engine uses Kymatio's internal scattering1d.filter_bank module and reports a warning when the installed version is not 0.3.x.
PyTorch / GPUNot required. The scattering gradient is implemented analytically in NumPy; no autograd or GPU is used.

Limitations

Cross-band modulation phase is not constrained. Time scattering can preserve modulation rate and depth within individual bands without preserving the phase relationship of that modulation across bands. A coherent full-band tremolo can therefore become less synchronized across the reconstructed bands.

Outputs

The script creates a new mono Sound. In Custom mode the normal name is:

<original-name>_ScatteringTexture

Named presets use concise output tags, for example:

<original-name>_SpectralSkeleton
<original-name>_TemporalTexture
<original-name>_ModulationGhost
<original-name>_StructureWithoutIdentity
<original-name>_SecondOrderReconstruction
<original-name>_RadicalTexture

The output preserves the source duration, start time, and sample rate. The original Sound remains unchanged.

Reproducibility: the Info window reports the seed actually used. Re-enter that seed with the same source and parameters to reproduce the same stochastic realization as far as the CPU implementation permits.

Applications

Texture-preserving re-synthesis

Use case: create a new realization of a sound while retaining broad spectrum, envelope behavior, or modulation statistics rather than its exact waveform.

Starting point: Custom, Medium scale, Medium detail, moderate order-1 weight and high order-2 weight.

Extract a spectral skeleton

Use case: retain broad timbral organisation but allow attacks, tremolo, and local temporal motion to diverge.

Starting point: Spectral Skeleton.

Preserve motion while changing identity

Use case: keep pulsation, articulation density, or modulation behavior while suppressing literal timbral similarity.

Starting point: Temporal Texture or Modulation Ghost.

Second-order experimental resynthesis

Use case: foreground what wavelet scattering adds beyond first-order CWT structure by making modulation constraints dominant.

Starting point: Second-Order Reconstruction.

Generate structurally related but strongly transformed material

Use case: create new material for electroacoustic composition in which the source's multiscale organisation remains perceptible while spectral and modulation coordinates are actively displaced.

Starting point: Radical Texture with increased Texture transformation and Spectral randomization.

Workflow: Tremolo source → modulation-preserving texture

Source: sustained instrumental sound with a clear amplitude modulation.
Settings: Temporal Texture or Second-Order Reconstruction; Preserve temporal modulation high; Medium or Broad texture scale.
Result: a new timbral realization whose individual scattering bands retain modulation-rate/depth characteristics even though exact waveform, pitch identity, and cross-band modulation phase are not preserved.

Workflow: Complex noisy sound → Structure Without Identity

Source: breath, scraping, multiphonic, field recording, or noisy instrumental texture.
Settings: Structure Without Identity; moderate-to-high Spectral randomization; adjust Texture transformation to taste.
Result: the output remains related to the source's broad multiscale behavior while becoming clearly detached from its literal acoustic identity.