Scattering Texture Generator — Wavelet-Scattering Texture Resynthesis
Generate a new waveform that preserves selected multiscale properties of a source — coarse energy, wavelet-band spectral structure, and second-order temporal modulation — while releasing exact waveform, phase, pitch identity, and fine timing. The process combines Kymatio Morlet scattering, structural target transformation, calibrated direct synthesis, and optional iterative refinement.
What this does
Scattering Texture Generator analyses a selected Sound with a one-dimensional Morlet wavelet-scattering representation and constructs a new waveform realization whose selected scattering statistics approach a source-derived or transformed target. Instead of copying the original waveform, the system preserves higher-level relationships such as spectral-band energy, coarse envelope shape, and multiscale modulation behavior.
Key Features:
- Orders 0, 1, and 2 — coarse energy envelope, wavelet-band structure, and second-order modulation.
- Normalized second-order matching — order 2 is matched as
S2 / S1, so modulation depth is not simply another measure of spectral energy. - 6 named creative presets + Custom — from Spectral Skeleton to Radical Texture.
- Target transformation — smoothing, spectral shifts, modulation-rate shifts, dynamic compression/expansion, sparsification, and structured perturbation.
- Spectral randomization — move carrier magnitude from the source spectrum toward a smoothed pitch-reduced spectral envelope.
- Calibrated direct synthesis — builds the waveform from the target before optional optimization.
- Optional L-BFGS refinement — Standard and High modes further reduce scattering-domain mismatch.
- Parallel segmented processing — long sounds are processed in independent overlapping segments and joined with equal-power crossfades.
- Reproducible variation — reuse a seed to recreate the same stochastic realization.
- Explanatory Praat Picture visualisation — source, multiscale structure, second-order modulation, target constraints, matching, and output.
- CPU-only and offline — no neural model, training, cloud service, API, GPU, or downloaded model weights.
CWT Scalogram — what multiscale components exist?
CWT Granular Resampler — use that structure to control another sound.
Scattering Texture Generator — generate a new waveform that preserves selected multiscale relationships.
Quick start
- Select exactly one Sound object in Praat.
- Run
Scattering_Texture_Generator.praat. - For a first run, leave Preset = Custom, Texture scale = Medium, and Frequency detail = Medium.
- Use the default preservation weights: Global envelope 40%, Spectral structure 60%, Temporal modulation 100%.
- Use Texture transformation to determine how strongly the scattering target itself is altered.
- Use Spectral randomization to progressively remove fine spectral / pitch identity from the carriers.
- Start with Preview for a fast direct-synthesis result. Use Standard or High when you want iterative refinement.
- Keep Draw visualization enabled to inspect source, target, achieved scattering structure, and loss behavior.
How it works
1. Praat exports a mono analysis signal
The selected Sound is preserved in the Object list. If it has more than one channel, Praat converts a temporary copy to mono for analysis and synthesis, applies a protective pre-scale only when required for 24-bit WAV export, and sends the temporary WAV to the Python backend.
2. Kymatio Morlet filters define the scattering representation
The Python engine uses Kymatio's one-dimensional Morlet filter bank. The forward scattering path is implemented in NumPy with Fourier-domain filtering, dyadic subsampling, modulus, and low-pass averaging. Orders 1 and 2 follow the Kymatio scattering structure; order 0 uses an envelope form appropriate for audio.
3. The source representation becomes a texture target
The source coefficients are moved into a log-domain target representation. A preset may apply its own structural operations, while Texture transformation scales a second group of target operations. These transformations act on the representation itself rather than functioning as a dry/wet control.
4. A new waveform is built from the target
For every first-order band, the engine creates a phase-randomized carrier. Its magnitude spectrum ranges from the source's detailed spectrum toward a smoothed pitch-reduced envelope according to Spectral randomization. Time-varying first-order envelopes and second-order Morlet-band modulation signals are then imposed on those carriers.
5. Calibration and optional refinement
The direct-synthesis stage repeatedly re-analyses the new waveform and corrects first-order levels and second-order modulation depths. Standard and High modes then use an analytic scattering gradient and L-BFGS to reduce a weighted log-domain mismatch.
Scattering orders — Musical interpretation
Order 0 — Global envelope
Representation: |x| * phi_J
Musical role: coarse amplitude shape and slow energy evolution.
Control: Preserve global envelope.
Order 1 — Spectral structure
Representation: averaged magnitudes of first-order Morlet bands.
Musical role: broad spectral distribution, band energy, and timbral trajectory.
Control: Preserve spectral structure.
Order 2 — Temporal modulation
Representation: modulation of the first-order band envelopes, matched in normalized S2/S1 form.
Musical role: pulsation, tremolo-like behavior, articulation density, roughness, and multiscale texture motion.
Control: Preserve temporal modulation.
6 Presets + Custom
Custom
Weights: user controls.
Transformation family: order-1 time smoothing, modulation-depth emphasis, and order-2 sparsification, scaled by Texture transformation.
Use: direct manual control of the three scattering-order weights.
Spectral Skeleton
Weights 0/1/2: 30 / 100 / 15.
Transformation: order-1 time smoothing and order-2 sparsification.
Use: retain broad timbral structure while loosening temporal modulation.
Temporal Texture
Weights 0/1/2: 30 / 35 / 100.
Preset operation: mild order-1 frequency blur. Transformation adds further blur and a downward spectral-envelope shift.
Use: foreground modulation behavior while reducing literal timbral identity.
Modulation Ghost
Weights 0/1/2: 60 / 10 / 100.
Preset operations: strong spectral blur and order-1 dynamic compression. Transformation increases modulation depth.
Use: create a temporal / textural shadow of the source.
Structure Without Identity
Weights 0/1/2: 50 / 50 / 50.
Preset operations: spectral blur and time smoothing. Transformation adds structured perturbation and an upward spectral-envelope shift.
Use: preserve broad multiscale behavior while weakening recognisable source identity.
Second-Order Reconstruction
Weights 0/1/2: 20 / 25 / 100.
Transformation: increased modulation depth.
Use: make second-order temporal structure the dominant reconstruction constraint.
Radical Texture
Weights 0/1/2: 40 / 60 / 100.
Transformation: modulation-rate shift, spectral-envelope shift, order-1 dynamic expansion, order-2 sparsification, and structured perturbation.
Use: produce a clearly transformed texture that still derives from the source's multiscale organization.
[preset, always]; operations scaled by the Transformation control are labelled [x Transformation].Controls
| Control | Default | Function |
|---|---|---|
| Preset | Custom | Selects Custom or one of six named texture strategies. Named presets set their own order-0 / order-1 / order-2 preservation weights. |
| Texture scale | Medium (~186 ms) | Sets the scattering averaging scale. Fine is approximately 46 ms; Medium 186 ms; Broad 743 ms. The exact J is derived from the working sample rate. |
| Frequency detail | Medium (Q = 8) | Controls first-order frequency resolution. Low uses Q=4, Medium Q=8, High Q=12; High also uses a denser second-order setting. |
| Preserve global envelope | 40% | Weights order 0, preserving coarse temporal energy behavior. |
| Preserve spectral structure | 60% | Weights order 1, preserving wavelet-band energy and timbral structure. |
| Preserve temporal modulation | 100% | Weights normalized order 2, preserving modulation depth and multiscale textural movement. |
| Texture transformation | 30% | Scales the preset's transformable operations. This changes the scattering target itself; it is not a dry/wet mix. |
| Spectral randomization | 50% | Blends carrier magnitude from the source's detailed spectrum toward a smoothed pitch-reduced spectral envelope. Carrier phase is randomized at every setting. |
| Reconstruction quality | Preview | Selects internal working-rate cap, segment size, and optional L-BFGS refinement. |
| Random seed | 1 | Positive values reproduce the same realization. 0 requests a new seed for each run. |
| Draw visualization | On | Creates the explanatory Praat Picture display after synthesis. |
| Play | On | Plays the generated Sound after it is imported and named. |
Reconstruction quality
| Mode | Internal rate cap | Calibration | L-BFGS refinement | Character |
|---|---|---|---|---|
| Preview | 16 kHz | 5 rounds | None | Fastest. Uses calibrated direct synthesis only. |
| Standard | 22.05 kHz | 5 rounds | Up to 15 iterations per segment | Balances speed and scattering-domain refinement. |
| High | 48 kHz | 5 rounds | Up to 150 iterations per segment | Highest working bandwidth and deepest iterative matching. |
Visualisation
When Draw visualization is enabled, the script produces a Praat Picture display organised as a five-stage explanation of the transformation:
1 — Source
Shows the mono analysis copy of the selected Sound and establishes the input waveform used by the backend.
2 — Multiscale analysis
Compares source, target, and achieved low-order scattering information. Grey represents the source, blue the transformed target, and red the new realization.
3 — What scattering adds
Displays order-2 modulation across carrier band and modulation rate, making explicit the layer that is not represented separately in a standard CWT scalogram.
4 — Texture constraints + iterative matching
Shows the active order weights and the evolution of the texture-matching loss, connecting the user controls to the reconstruction process.
5 — New realization
Shows the resulting waveform together with a summary of scattering geometry, target operations, sample-rate behavior, segmentation, refinement, seed, and warnings.
Technical behavior
- Processes exactly one selected Sound and never modifies the original object.
- Converts multichannel input to a temporary mono analysis copy; the generated result is mono.
- Exports a temporary 24-bit WAV, with protective pre-scaling only when necessary.
- Uses Kymatio's Morlet filter-bank definitions but performs scattering and its analytic adjoint in NumPy.
- Order 0 uses
|x| * phi_Jbecause the ordinary low-passed audio waveform is near zero for DC-free audio. - Order 2 is matched in normalized
S2/S1form so temporal-modulation preservation does not simply re-impose spectral-band level. - Coefficient-family losses are computed in the log domain and normalized by coefficient count.
- Direct synthesis creates first-order carriers and second-order modulation sources using the same Morlet families used by the analysis.
- Long sounds are segmented with context margins and reassembled with equal-power crossfades.
- Segment jobs can run in parallel worker processes; worker FFTs use a single thread to avoid CPU oversubscription.
- If multiprocessing cannot start, the backend falls back to a single process.
- Silent input returns silence of identical duration.
- Output sample count and sample rate are checked against the source before the Sound is accepted.
- Output RMS is matched to the source and peak protection attenuates only when required.
- Temporary files are cleaned after processing.
- The Info window reports the seed, order weights,
J, averaging scale, Q values, path counts, modulation-rate range, working rate, band limit, segmentation, calibration, refinement, target operations, loss reduction, and warnings.
Requirements & installation
numpy, scipy, and kymatio.Install with:
python -m pip install numpy scipy kymatio
| Component | Requirement |
|---|---|
| Praat | Praat 6.1+; the script notes testing with 6.1.38, 6.4.06, and 7.0. Praat 7 may ask for full trust because the tool writes temporary files and launches Python. |
| Python | Python 3. The frontend uses the library's OS-specific Python discovery convention. |
| Python backend | Place scattering_texture_engine.py in plugin_AudioTools/py/. The script also accepts the engine next to the Praat script as a fallback. |
| Kymatio | The engine uses Kymatio's internal scattering1d.filter_bank module and reports a warning when the installed version is not 0.3.x. |
| PyTorch / GPU | Not required. The scattering gradient is implemented analytically in NumPy; no autograd or GPU is used. |
Limitations
- Not an exact reconstruction: scattering is used as a texture constraint, so the output is intentionally a new realization rather than a waveform copy.
- Mono output: multichannel source material is mixed to mono before scattering analysis and resynthesis.
- Pitch identity is not explicitly preserved: the method can retain broad spectral structure while exact harmonic phase and fine pitch identity are deliberately released.
- Spectral randomization does not disable phase randomization: the control changes the carrier magnitude spectrum; random carrier phase is used throughout the direct-synthesis process.
- Preset operations can be active at Transformation 0: some named presets intentionally contain always-on structural operations.
- Quality changes internal bandwidth: Preview, Standard, and High use different working-rate caps in addition to different refinement depths.
- CPU cost can be substantial: fine frequency detail, broad texture scale, long sounds, and High refinement require more computation and memory.
- Minimum source duration: the Praat frontend rejects Sounds shorter than 20 ms.
Outputs
The script creates a new mono Sound. In Custom mode the normal name is:
<original-name>_ScatteringTexture
Named presets use concise output tags, for example:
<original-name>_SpectralSkeleton <original-name>_TemporalTexture <original-name>_ModulationGhost <original-name>_StructureWithoutIdentity <original-name>_SecondOrderReconstruction <original-name>_RadicalTexture
The output preserves the source duration, start time, and sample rate. The original Sound remains unchanged.
Applications
Texture-preserving re-synthesis
Use case: create a new realization of a sound while retaining broad spectrum, envelope behavior, or modulation statistics rather than its exact waveform.
Starting point: Custom, Medium scale, Medium detail, moderate order-1 weight and high order-2 weight.
Extract a spectral skeleton
Use case: retain broad timbral organisation but allow attacks, tremolo, and local temporal motion to diverge.
Starting point: Spectral Skeleton.
Preserve motion while changing identity
Use case: keep pulsation, articulation density, or modulation behavior while suppressing literal timbral similarity.
Starting point: Temporal Texture or Modulation Ghost.
Second-order experimental resynthesis
Use case: foreground what wavelet scattering adds beyond first-order CWT structure by making modulation constraints dominant.
Starting point: Second-Order Reconstruction.
Generate structurally related but strongly transformed material
Use case: create new material for electroacoustic composition in which the source's multiscale organisation remains perceptible while spectral and modulation coordinates are actively displaced.
Starting point: Radical Texture with increased Texture transformation and Spectral randomization.
Workflow: Tremolo source → modulation-preserving texture
Source: sustained instrumental sound with a clear amplitude modulation.
Settings: Temporal Texture or Second-Order Reconstruction; Preserve temporal modulation high; Medium or Broad texture scale.
Result: a new timbral realization whose individual scattering bands retain modulation-rate/depth characteristics even though exact waveform, pitch identity, and cross-band modulation phase are not preserved.
Workflow: Complex noisy sound → Structure Without Identity
Source: breath, scraping, multiphonic, field recording, or noisy instrumental texture.
Settings: Structure Without Identity; moderate-to-high Spectral randomization; adjust Texture transformation to taste.
Result: the output remains related to the source's broad multiscale behavior while becoming clearly detached from its literal acoustic identity.