Phonetic Tremolo/Glitch Effect — User Guide
An acoustic-class-driven processor that analyzes a selected Sound, groups the recording into vowel-like, fricative-like, silence, and other regions, and applies a different time-domain treatment to each class.
What this does
Phonetic Tremolo/Glitch Effect analyzes the recording as a sequence of short acoustic frames and assigns each frame to one of four practical processing classes: vowel-like, fricative-like, silence, or other. Adjacent frames with the same class are merged into larger regions, and each region is processed according to its class.
Vowel-like regions receive tremolo, fricative-like regions receive a short delayed replacement, silence is attenuated, and other material remains unchanged. A short transition ramp can blend each processed region back into the untouched source at its boundaries, reducing abrupt changes between neighboring classifications.
Processing pipeline
1. Select the analysis signal
Mono input is analyzed directly. For multichannel input, the script first creates a mono fold-down. It also measures the peak of every original channel. If the fold-down peak is less than 10% of the strongest channel peak, the fold-down is treated as potentially cancelled by opposing channel content and the strongest individual channel is used for analysis instead.
2. Extract acoustic features
Intensity is always calculated because silence is always one of the processing classes. When the active effect or visualization requires detailed classification, the script also creates:
- Pitch — time step = Frame step, floor 75 Hz, ceiling = min(600 Hz, Nyquist − 50 Hz).
- Formant (Burg) — five formants, time step = Frame step, 25 ms window, 50 Hz pre-emphasis setting, with the formant ceiling limited by Nyquist.
- Harmonicity (cc) — time step = Frame step, minimum pitch 75 Hz.
If tremolo and fricative delay are both disabled and visualization is also disabled, the processor uses an intensity-only fast path: it only distinguishes silence from other material because the finer classes cannot change the resulting audio.
3. Classify frames
The Sound is divided into successive frames of Frame step duration. One set of acoustic measurements is read at the midpoint of each frame. Classification is completed for the whole recording before any audio is modified.
4. Merge neighboring classes
Consecutive frames with the same class are merged into one processing region. Effects are therefore applied to contiguous regions rather than independently resetting on every analysis frame.
5. Process each region against the untouched reference
The script keeps an unchanged reference copy of the original Sound. Every class effect reads from that reference, so one processed region does not become the source for the next region.
6. Safety ceiling
After processing, the peak is measured. If active processing produced a peak above Safety peak, the complete result is scaled down to that ceiling. Material already below the ceiling is not raised.
Acoustic classification
Silence
Silence is tested first. Undefined Intensity values are treated as extremely low level.
Vowel-like
A non-silent frame is classified as vowel-like when it is voiced, sufficiently harmonic, and has a plausible three-formant structure.
Formant-structure validation
F1, F2, F3 and their bandwidths must all be defined. The implementation then requires:
- F1 > Vowel F1 minimum.
- F2 > F1 and F3 > F2.
- F2 − F1 ≥ 120 Hz.
- F3 − F2 ≥ 180 Hz.
- F3 − F1 ≥ max(450 Hz, 1.25 × F0 when voiced).
- F3 below the resolved formant-analysis ceiling.
- All three bandwidths positive and sufficiently narrow relative to their formant frequencies.
This extra structure test prevents the vowel-like decision from relying only on a single F1 value.
Fricative-like
This branch is evaluated only after the frame has already passed the silence test. It therefore selects audible, unvoiced, noise-like frames.
Other
Every remaining non-silent frame is assigned to Other. This material remains unchanged by the class-specific audio stage.
Class-based effects
Vowel-like → Tremolo
The vowel-like region is multiplied by a unipolar sinusoidal attenuation curve. At the full body of a region:
At 100% Wet, the gain therefore moves between 1 and 1 − Tremolo_depth. A depth of 1.0 can reach zero at the deepest part of the tremolo cycle.
Fricative-like → Delayed replacement
The fricative effect crossfades from the original sample to an earlier sample from the same channel:
This is a short backward-looking delay, not a forward read. At 100% Wet and in the body of the region, the fricative-like material is replaced by the delayed source. At intermediate Wet values it is mixed with the undelayed source.
Silence → Attenuation
At 100% Wet and away from the transition edges, a Silence Gain of 0.05 leaves 5% of the source amplitude; a value of 0 produces full gating.
Other → Unchanged
Other regions are copied from the original reference without amplification or additional processing.
Region transitions
Transition ms creates a linear fade of the class effect at both edges of each merged region. The effective transition length is:
The effect is 0 at the exact region boundary, rises to full strength through the entry transition, remains full in the body, and falls back to 0 through the exit transition. Setting Transition to 0 ms disables this edge blending.
Global Dry/Wet
Dry/Wet scales the strength of every class-specific transformation. At 0% Wet the script uses an exact bypass fast path: it simply copies the source, skips acoustic analysis and Safety Peak scaling, and labels the analysis source as bypassed.
Presets
| Preset | Tremolo | Fricative delay | Silence settings |
|---|---|---|---|
| Custom | 8 Hz, depth .70 by default | 15 ms | threshold 45 dB, gain .05 |
| Subtle Vocal Texture | 4 Hz, depth .30 | 5 ms | threshold 40 dB, gain .10 |
| Hard Robot Glitch | 12 Hz, depth .90 | 30 ms | threshold 50 dB, gain .02 |
| Broken Radio (High Speed) | 25 Hz, depth .80 | 10 ms | threshold 45 dB, gain .03 |
| Fricative Smear (Long Delay) | 6 Hz, depth .20 | 80 ms | gain .05; Fricative HNR max becomes 5 dB |
| Deep Vowel Tremolo | 15 Hz, depth 1.00 | 0 ms | gain .05 |
| Clean Gated (Silence Removal) | disabled | disabled | threshold 60 dB, gain 0 |
Parameters
Main form
| Parameter | Default | What it controls | Runtime range |
|---|---|---|---|
| Tremolo rate Hz | 8.0 | Vowel-like tremolo rate. | Minimum 0 Hz. |
| Tremolo depth | 0.7 | Maximum attenuation depth on vowel-like regions. | 0…1. |
| Fricative delay seconds | 0.015 | Backward-looking delay used on fricative-like regions. | 0…Sound duration. |
| Silence gain | 0.05 | Residual gain in silence regions at full Wet. | 0…1. |
| Transition ms | 2.0 | Linear effect ramp at both edges of every merged region. | 0…50 ms; also limited to half the region duration. |
| Dry/Wet percent | 100 | Global amount of all class-specific processing. | 0…100%. |
| Advanced settings | off | Opens the analysis/threshold/safety dialog. | — |
| Draw visualization | on | Draws the AudioTools figure. | Does not change the audio. |
| Play result | on | Plays the finished Sound. | Does not change the audio. |
Advanced settings
| Parameter | Default | Meaning | Runtime handling |
|---|---|---|---|
| Frame step seconds | 0.01 | Classification-frame duration and analysis time step. | Clamped to .002….1 s. |
| Max formant Hz | 5500 | Requested Burg formant ceiling. | Resolved to min(user value, Nyquist − 50 Hz). |
| Vowel HNR threshold | 5 dB | Minimum Harmonicity for vowel-like classification. | Used directly. |
| Vowel F1 minimum Hz | 300 | Minimum F1 inside the formant-structure validator. | Used directly. |
| Fricative HNR max | 3 dB | Maximum Harmonicity for fricative-like classification. | Used directly. |
| Silence intensity threshold | 45 dB | Frames below this value become Silence before other tests. | Used directly; some presets override it. |
| Safety peak | 0.99 | Maximum allowed peak after active processing. | Clamped to 0…1; 0 disables scaling; attenuation only. |
Channels & output
| Property | Behavior |
|---|---|
| Analysis | Mono input directly; otherwise mono fold-down with strongest-channel fallback if the fold-down is strongly cancelled. |
| Effect control | One shared acoustic classification timeline controls all output channels. |
| Audio processing | Applied to every original channel over the same classified regions; the output is created from the original multichannel Sound. |
| Channel count | Preserved. |
| Sample rate | Preserved. |
| Duration / time domain | Preserved. |
| Randomness | None. The same source and settings produce the same classification and processing. |
| Safety | Only active results above Safety Peak are attenuated. |
Visualization
When Draw visualization is enabled, v0.4 creates an AudioTools figure with:
Input
Original waveform.
Output
Final processed waveform.
Acoustic-class timeline
Color-coded vowel-like, fricative-like, silence, and other classifications across source time.
Legend & counts
Full-frame counts for the four classes.
Long-file timeline display
The classification itself is calculated for every frame. For drawing, the timeline displays at most 500 representative frame positions. The class counts and processing still use the complete frame sequence.
Summary strip
The bottom strip reports tremolo rate/depth, fricative delay, Silence Gain, Transition, Wet percentage, silence threshold, chosen analysis source, duration, sample rate, and channel count.