Entropy Smart De-Esser — User Guide
Split-band de-essing that combines high-frequency level and HF/full-band ratio with spectral-entropy confidence, then reduces only the selected high-frequency band.
What this does
Entropy Smart De-Esser reduces sibilant, noise-like high-frequency energy while leaving the lower body of the signal at its original level. It is a split-band processor: the selected HF band is extracted, a time-varying gain is calculated, and only that HF band is attenuated.
output = residual + HF band × gain
Detection uses three complementary cues. The HF band must first be loud enough in absolute level; its amplitude must also be large relative to the full-band signal; and spectral entropy inside that same HF band then determines how strongly the ratio-derived reduction should be trusted. Noise-like, broadly distributed HF receives more of the requested reduction than tonal or strongly peaked HF.
Quick start
- Select exactly one Sound object in Praat.
- Run
ENTROPY_SMART_DE-ESSER_v0.8.praat. - Choose a preset, or use Custom.
- For normal full-band material, start with the default HF band of 4–8 kHz.
- Use Threshold to decide how strong the HF/full-band ratio must be before reduction develops.
- Use Max_reduction_db to set the maximum possible attenuation of the HF band.
- Leave Listen_to_removed off for the processed signal; turn it on to audition only the HF content being removed.
Presets
Presets override the listed detector and envelope controls. Other settings, including Dry/Wet, output level, monitoring and visualization, remain available from the form.
| Preset | HF band | Threshold | Max reduction | Attack / Release |
|---|---|---|---|---|
| Custom | 4–8 kHz | 0.40 | 12 dB | 5 / 50 ms |
| Light De-Essing | current form values | 0.50 | 6 dB | 5 / 60 ms |
| Medium De-Essing | current form values | 0.40 | 10 dB | 5 / 50 ms |
| Heavy De-Essing | current form values | 0.30 | 15 dB | 3 / 40 ms |
| Aggressive | current form values | 0.25 | 18 dB | 2 / 30 ms |
| Gentle Touch | current form values | 0.55 | 4 dB | 8 / 80 ms |
| Telephony band | 2.5–3.7 kHz | 0.40 | 10 dB | 5 / 50 ms |
The Telephony preset is intended for narrowband material where the normal 4–8 kHz region is unavailable or partly above Nyquist.
Controls
| Control | Default | What it changes |
|---|---|---|
| Hf_low_hz | 4000 Hz | Lower edge of the band used both for HF-ratio detection and for split-band attenuation. |
| Hf_high_hz | 8000 Hz | Upper edge. It is automatically limited to 95% of Nyquist; the remaining band must be at least 200 Hz wide. |
| Threshold | 0.40 | Centre of the HF/full-band ratio knee. Lower values make the detector more sensitive. |
| Max_reduction_db | 12 dB | Lower limit of the HF gain. It does not attenuate low frequencies. |
| Attack_ms | 5 ms | Smoothing applied when the HF gain is falling toward greater reduction. |
| Release_ms | 50 ms | Smoothing applied when the HF gain returns toward unity. |
| Dry_wet_mix | 1.0 | Blends the de-essed signal with the original. It is used only for the normal processed output, not removed-only monitoring. |
| Output_level_mode | Natural level | Natural level, Match input RMS, Safety ceiling, or Peak normalize. |
| Ceiling_peak | 0.95 | Target used by Safety ceiling and Peak normalize. |
| Listen_to_removed | Off | Outputs only the removed HF component instead of the de-essed signal. |
| Show_spectrogram | On | Includes the result spectrogram in the Picture visualization. |
| Draw_visualization | On | Draws the diagnostic page. |
| Play_after_processing | On | Plays the result after processing. |
How detection works
1. HF level and HF/full-band ratio
For each detector channel, the script measures full-band intensity and the intensity of a Hann-pass HF band. It converts the two dB values back to linear amplitude and forms:
The ratio is limited to 0–1. A frame can drive the detector only when the HF level also clears the internal absolute gate of −55 dBFS. This prevents very quiet broadband noise from being interpreted as meaningful sibilance merely because its HF/full-band ratio is large.
2. Spectral entropy inside the HF band
For candidate frames, a 20 ms Gaussian spectrogram is sampled only between Hf_low_hz and Hf_high_hz. The power values in that band are normalized to a probability distribution, and normalized Shannon entropy is computed:
H = −Σ p[k] ln p[k] / ln(N)
Low entropy means the HF energy is concentrated in relatively few spectral bins; high entropy means it is distributed more broadly and is more noise-like. Entropy therefore describes spectral shape, while the HF gate and ratio describe amount and relative prominence.
3. Soft entropy confidence
The internal entropy range 0.55–0.90 is mapped smoothly to confidence 0–1. With the default entropy weight of 0.70, low-entropy HF retains 30% of the reduction requested by the ratio detector, while high-entropy HF can receive 100%.
requested reduction = ratio reduction × entropy scale
This is intentionally not a binary gate: tonal HF can still be controlled when the HF-ratio detector strongly identifies it, but it is protected from the full attenuation normally reserved for noise-like sibilance.
4. Ratio knee and gain smoothing
The ratio detector uses an internal knee width of 0.12 centred on Threshold. Inside the knee, the transition is quadratic and continuous in slope. The entropy-scaled target gain is then smoothed using separate Attack and Release times.
Audio processing
Signal path
Input → detector analysis → HF ratio + absolute gate → HF-band entropy confidence → soft-knee target gain → attack/release → split-band HF attenuation → optional Dry/Wet → optional output-level stage.
The same selected frequency range is extracted with Filter (pass Hann band) for processing. For the normal output:
At gain = 1 the equation reduces exactly to the input. The low-frequency residual is not multiplied by the de-essing gain, so a sibilant event does not intentionally duck the fundamental, low formants or other content outside the selected HF band.
No random process is used; identical input and settings produce deterministic processing.
Channels & timing
Every input channel is processed and retained. By default, detection examines all channels and, frame by frame, uses the channel with the strongest valid HF evidence to drive one shared gain curve. Applying the same gain trajectory to every channel preserves inter-channel level relationships while still allowing a sibilant event on any channel to trigger the processor.
The Sound is processed on a temporary copy shifted to time 0, and the original start time is restored to the result. Sample rate, duration and number of channels are preserved.
The script uses a conservative short-file guard and requires at least about 48 ms of input before running the detector.
Output & monitoring
Normal processed output
With Listen_to_removed = Off, the output is named:
Removed-only monitoring
With Listen_to_removed = On, the output contains only what the gain curve removed from the selected HF band:
name = <original>_sibilants
Dry/Wet is not applied to this monitoring output.
Output level modes
| Mode | Behavior |
|---|---|
| Natural level | No final level scaling. |
| Match input RMS | Scales the complete output to the input Sound's RMS. |
| Safety ceiling | Attenuates only when the output peak exceeds Ceiling_peak. |
| Peak normalize | Always scales a non-silent output so its peak equals Ceiling_peak. |
If the stored result peaks above 1.0, the script warns about possible clipping in integer PCM. For playback only, it creates a temporary copy scaled to 0.95; the stored Sound itself is not changed by that playback safeguard.
Visualization
The optional 8 × 8 Picture page contains:
- Waveform + gain reduction: a real input channel is shown in grey with the reduction envelope in red.
- Detector panel: HF/full ratio in blue, HF-band entropy in orange, the threshold/knee region in pink, and frames rejected by the absolute HF gate in green.
- Zoom comparison: the first up to 2 seconds of one real channel, original in grey and output in blue.
- Result spectrogram: shown when
Show_spectrogramis enabled, up to 8 kHz or Nyquist. - Summary: preset, HF band, threshold and knee, entropy range, maximum reduction, attack/release, number of reduced frames, gated percentage, Dry/Wet, peaks and output-level action.
The waveform display deliberately uses a real channel rather than a mono fold, so anti-phase multichannel material cannot disappear from the diagnostic plot through cancellation.
Further Reading
- Zölzer, U. (ed.) (2011). DAFX: Digital Audio Effects, 2nd ed., Wiley, Chapter 4: Nonlinear Processing — includes dynamic-range control and de-essing. DAFX chapter page.
- Shannon, C. E. (1948). “A Mathematical Theory of Communication.” Bell System Technical Journal, 27, 379–423 and 623–656. Reprint.
- Strevens, P. (1960). “Spectra of Fricative Noise in Human Speech.” Language and Speech, 3(1), 32–49. DOI: 10.1177/002383096000300105. SAGE.
- Boersma, P. & Weenink, D. Praat Manual: Sound: To Intensity... — intensity window and automatic time-step behavior. Praat manual.