Phonetic Tremolo/Glitch Effect — User Guide

An acoustic-class-driven processor that analyzes a selected Sound, groups the recording into vowel-like, fricative-like, silence, and other regions, and applies a different time-domain treatment to each class.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 0.4 (2026) License: MIT License Repo: https://github.com/ShaiCohen-ops/Praat-plugin_AudioTools
Contents:

What this does

Phonetic Tremolo/Glitch Effect analyzes the recording as a sequence of short acoustic frames and assigns each frame to one of four practical processing classes: vowel-like, fricative-like, silence, or other. Adjacent frames with the same class are merged into larger regions, and each region is processed according to its class.

Vowel-like regions receive tremolo, fricative-like regions receive a short delayed replacement, silence is attenuated, and other material remains unchanged. A short transition ramp can blend each processed region back into the untouched source at its boundaries, reducing abrupt changes between neighboring classifications.

The classification is an acoustic control layer. It is derived from Intensity, Pitch, Harmonicity, and formant-structure tests and is used to decide where each effect is applied. The labels therefore describe the acoustic behavior used by the processor: vowel-like, fricative-like, silence, and other.

Processing pipeline

1. Select the analysis signal

Mono input is analyzed directly. For multichannel input, the script first creates a mono fold-down. It also measures the peak of every original channel. If the fold-down peak is less than 10% of the strongest channel peak, the fold-down is treated as potentially cancelled by opposing channel content and the strongest individual channel is used for analysis instead.

2. Extract acoustic features

Intensity is always calculated because silence is always one of the processing classes. When the active effect or visualization requires detailed classification, the script also creates:

If tremolo and fricative delay are both disabled and visualization is also disabled, the processor uses an intensity-only fast path: it only distinguishes silence from other material because the finer classes cannot change the resulting audio.

3. Classify frames

The Sound is divided into successive frames of Frame step duration. One set of acoustic measurements is read at the midpoint of each frame. Classification is completed for the whole recording before any audio is modified.

4. Merge neighboring classes

Consecutive frames with the same class are merged into one processing region. Effects are therefore applied to contiguous regions rather than independently resetting on every analysis frame.

5. Process each region against the untouched reference

The script keeps an unchanged reference copy of the original Sound. Every class effect reads from that reference, so one processed region does not become the source for the next region.

6. Safety ceiling

After processing, the peak is measured. If active processing produced a peak above Safety peak, the complete result is scaled down to that ceiling. Material already below the ceiling is not raised.

Acoustic classification

Silence

if Intensity < Silence_intensity_threshold class = Silence

Silence is tested first. Undefined Intensity values are treated as extremely low level.

Vowel-like

A non-silent frame is classified as vowel-like when it is voiced, sufficiently harmonic, and has a plausible three-formant structure.

HNR > Vowel_hnr_threshold F0 > 0 AND formant structure passes validation

Formant-structure validation

F1, F2, F3 and their bandwidths must all be defined. The implementation then requires:

This extra structure test prevents the vowel-like decision from relying only on a single F1 value.

Fricative-like

HNR < Fricative_hnr_max AND F0 = 0

This branch is evaluated only after the frame has already passed the silence test. It therefore selects audible, unvoiced, noise-like frames.

Other

Every remaining non-silent frame is assigned to Other. This material remains unchanged by the class-specific audio stage.

The Info window reports the number of vowel-like, fricative-like, silence, and other frames, plus the number of frames that passed the formant-structure validation.

Class-based effects

Vowel-like → Tremolo

The vowel-like region is multiplied by a unipolar sinusoidal attenuation curve. At the full body of a region:

LFO(t) = 0.5 × [1 + sin(2π × Tremolo_rate × local_time)] gain(t) = 1 − Wet × Tremolo_depth × LFO(t)

At 100% Wet, the gain therefore moves between 1 and 1 − Tremolo_depth. A depth of 1.0 can reach zero at the deepest part of the tremolo cycle.

Fricative-like → Delayed replacement

The fricative effect crossfades from the original sample to an earlier sample from the same channel:

processed(t) = original(t) × [1 − Wet × edge] + original(t − Fricative_delay) × Wet × edge

This is a short backward-looking delay, not a forward read. At 100% Wet and in the body of the region, the fricative-like material is replaced by the delayed source. At intermediate Wet values it is mixed with the undelayed source.

Silence → Attenuation

gain = 1 − Wet × edge × (1 − Silence_gain)

At 100% Wet and away from the transition edges, a Silence Gain of 0.05 leaves 5% of the source amplitude; a value of 0 produces full gating.

Other → Unchanged

Other regions are copied from the original reference without amplification or additional processing.

Region transitions

Transition ms creates a linear fade of the class effect at both edges of each merged region. The effective transition length is:

fade = min(user transition, 0.5 × region duration)

The effect is 0 at the exact region boundary, rises to full strength through the entry transition, remains full in the body, and falls back to 0 through the exit transition. Setting Transition to 0 ms disables this edge blending.

Global Dry/Wet

Dry/Wet scales the strength of every class-specific transformation. At 0% Wet the script uses an exact bypass fast path: it simply copies the source, skips acoustic analysis and Safety Peak scaling, and labels the analysis source as bypassed.

Presets

Preset Tremolo Fricative delay Silence settings
Custom 8 Hz, depth .70 by default 15 ms threshold 45 dB, gain .05
Subtle Vocal Texture 4 Hz, depth .30 5 ms threshold 40 dB, gain .10
Hard Robot Glitch 12 Hz, depth .90 30 ms threshold 50 dB, gain .02
Broken Radio (High Speed) 25 Hz, depth .80 10 ms threshold 45 dB, gain .03
Fricative Smear (Long Delay) 6 Hz, depth .20 80 ms gain .05; Fricative HNR max becomes 5 dB
Deep Vowel Tremolo 15 Hz, depth 1.00 0 ms gain .05
Clean Gated (Silence Removal) disabled disabled threshold 60 dB, gain 0
Presets override the musical controls listed above and, where shown, the relevant classifier threshold. Transition ms and Dry/Wet remain under the user's direct control.

Parameters

Main form

ParameterDefaultWhat it controlsRuntime range
Tremolo rate Hz8.0Vowel-like tremolo rate.Minimum 0 Hz.
Tremolo depth0.7Maximum attenuation depth on vowel-like regions.0…1.
Fricative delay seconds0.015Backward-looking delay used on fricative-like regions.0…Sound duration.
Silence gain0.05Residual gain in silence regions at full Wet.0…1.
Transition ms2.0Linear effect ramp at both edges of every merged region.0…50 ms; also limited to half the region duration.
Dry/Wet percent100Global amount of all class-specific processing.0…100%.
Advanced settingsoffOpens the analysis/threshold/safety dialog.—
Draw visualizationonDraws the AudioTools figure.Does not change the audio.
Play resultonPlays the finished Sound.Does not change the audio.

Advanced settings

ParameterDefaultMeaningRuntime handling
Frame step seconds0.01Classification-frame duration and analysis time step.Clamped to .002….1 s.
Max formant Hz5500Requested Burg formant ceiling.Resolved to min(user value, Nyquist − 50 Hz).
Vowel HNR threshold5 dBMinimum Harmonicity for vowel-like classification.Used directly.
Vowel F1 minimum Hz300Minimum F1 inside the formant-structure validator.Used directly.
Fricative HNR max3 dBMaximum Harmonicity for fricative-like classification.Used directly.
Silence intensity threshold45 dBFrames below this value become Silence before other tests.Used directly; some presets override it.
Safety peak0.99Maximum allowed peak after active processing.Clamped to 0…1; 0 disables scaling; attenuation only.
Detailed formant-based classification requires enough sample-rate bandwidth. If the resolved formant ceiling falls below 1200 Hz while detailed analysis is required, the script stops instead of performing a low-bandwidth classification.

Channels & output

PropertyBehavior
AnalysisMono input directly; otherwise mono fold-down with strongest-channel fallback if the fold-down is strongly cancelled.
Effect controlOne shared acoustic classification timeline controls all output channels.
Audio processingApplied to every original channel over the same classified regions; the output is created from the original multichannel Sound.
Channel countPreserved.
Sample ratePreserved.
Duration / time domainPreserved.
RandomnessNone. The same source and settings produce the same classification and processing.
SafetyOnly active results above Safety Peak are attenuated.

Visualization

When Draw visualization is enabled, v0.4 creates an AudioTools figure with:

Input

Original waveform.

Output

Final processed waveform.

Acoustic-class timeline

Color-coded vowel-like, fricative-like, silence, and other classifications across source time.

Legend & counts

Full-frame counts for the four classes.

Long-file timeline display

The classification itself is calculated for every frame. For drawing, the timeline displays at most 500 representative frame positions. The class counts and processing still use the complete frame sequence.

Summary strip

The bottom strip reports tremolo rate/depth, fricative delay, Silence Gain, Transition, Wet percentage, silence threshold, chosen analysis source, duration, sample rate, and channel count.

The visualization shows the classifier that controls the processing. The class timeline is derived from the analysis stage; it is not inferred from the processed output.