Entropy Smart De-Esser — User Guide

Split-band de-essing that combines high-frequency level and HF/full-band ratio with spectral-entropy confidence, then reduces only the selected high-frequency band.

Author: Shai Cohen Affiliation: Department of Music, Bar-Ilan University, Israel Version: 0.8 (2026) License: MIT License Repo: Praat AudioTools
Contents:

What this does

Entropy Smart De-Esser reduces sibilant, noise-like high-frequency energy while leaving the lower body of the signal at its original level. It is a split-band processor: the selected HF band is extracted, a time-varying gain is calculated, and only that HF band is attenuated.

residual = input − HF band
output = residual + HF band × gain

Detection uses three complementary cues. The HF band must first be loud enough in absolute level; its amplitude must also be large relative to the full-band signal; and spectral entropy inside that same HF band then determines how strongly the ratio-derived reduction should be trusted. Noise-like, broadly distributed HF receives more of the requested reduction than tonal or strongly peaked HF.

Important: entropy is a confidence modifier, not a stand-alone trigger. A high entropy value by itself does not cause attenuation. The HF-level gate and HF/full-band ratio establish the candidate first.

Quick start

  1. Select exactly one Sound object in Praat.
  2. Run ENTROPY_SMART_DE-ESSER_v0.8.praat.
  3. Choose a preset, or use Custom.
  4. For normal full-band material, start with the default HF band of 4–8 kHz.
  5. Use Threshold to decide how strong the HF/full-band ratio must be before reduction develops.
  6. Use Max_reduction_db to set the maximum possible attenuation of the HF band.
  7. Leave Listen_to_removed off for the processed signal; turn it on to audition only the HF content being removed.
Practical check: the removed-signal mode should contain mainly troublesome sibilant/high-frequency material. If vowels, body or low-frequency content are prominent, the selected HF band or detector settings are too broad or aggressive for that source.

Presets

Presets override the listed detector and envelope controls. Other settings, including Dry/Wet, output level, monitoring and visualization, remain available from the form.

PresetHF bandThresholdMax reductionAttack / Release
Custom4–8 kHz0.4012 dB5 / 50 ms
Light De-Essingcurrent form values0.506 dB5 / 60 ms
Medium De-Essingcurrent form values0.4010 dB5 / 50 ms
Heavy De-Essingcurrent form values0.3015 dB3 / 40 ms
Aggressivecurrent form values0.2518 dB2 / 30 ms
Gentle Touchcurrent form values0.554 dB8 / 80 ms
Telephony band2.5–3.7 kHz0.4010 dB5 / 50 ms

The Telephony preset is intended for narrowband material where the normal 4–8 kHz region is unavailable or partly above Nyquist.

Controls

ControlDefaultWhat it changes
Hf_low_hz4000 HzLower edge of the band used both for HF-ratio detection and for split-band attenuation.
Hf_high_hz8000 HzUpper edge. It is automatically limited to 95% of Nyquist; the remaining band must be at least 200 Hz wide.
Threshold0.40Centre of the HF/full-band ratio knee. Lower values make the detector more sensitive.
Max_reduction_db12 dBLower limit of the HF gain. It does not attenuate low frequencies.
Attack_ms5 msSmoothing applied when the HF gain is falling toward greater reduction.
Release_ms50 msSmoothing applied when the HF gain returns toward unity.
Dry_wet_mix1.0Blends the de-essed signal with the original. It is used only for the normal processed output, not removed-only monitoring.
Output_level_modeNatural levelNatural level, Match input RMS, Safety ceiling, or Peak normalize.
Ceiling_peak0.95Target used by Safety ceiling and Peak normalize.
Listen_to_removedOffOutputs only the removed HF component instead of the de-essed signal.
Show_spectrogramOnIncludes the result spectrogram in the Picture visualization.
Draw_visualizationOnDraws the diagnostic page.
Play_after_processingOnPlays the result after processing.

How detection works

1. HF level and HF/full-band ratio

For each detector channel, the script measures full-band intensity and the intensity of a Hann-pass HF band. It converts the two dB values back to linear amplitude and forms:

HF ratio = HF amplitude / full-band amplitude

The ratio is limited to 0–1. A frame can drive the detector only when the HF level also clears the internal absolute gate of −55 dBFS. This prevents very quiet broadband noise from being interpreted as meaningful sibilance merely because its HF/full-band ratio is large.

2. Spectral entropy inside the HF band

For candidate frames, a 20 ms Gaussian spectrogram is sampled only between Hf_low_hz and Hf_high_hz. The power values in that band are normalized to a probability distribution, and normalized Shannon entropy is computed:

p[k] = P[k] / ΣP[k]
H = −Σ p[k] ln p[k] / ln(N)

Low entropy means the HF energy is concentrated in relatively few spectral bins; high entropy means it is distributed more broadly and is more noise-like. Entropy therefore describes spectral shape, while the HF gate and ratio describe amount and relative prominence.

3. Soft entropy confidence

The internal entropy range 0.55–0.90 is mapped smoothly to confidence 0–1. With the default entropy weight of 0.70, low-entropy HF retains 30% of the reduction requested by the ratio detector, while high-entropy HF can receive 100%.

entropy scale = (1 − weight) + weight × confidence
requested reduction = ratio reduction × entropy scale

This is intentionally not a binary gate: tonal HF can still be controlled when the HF-ratio detector strongly identifies it, but it is protected from the full attenuation normally reserved for noise-like sibilance.

4. Ratio knee and gain smoothing

The ratio detector uses an internal knee width of 0.12 centred on Threshold. Inside the knee, the transition is quadratic and continuous in slope. The entropy-scaled target gain is then smoothed using separate Attack and Release times.

Audio processing

Signal path

Input → detector analysis → HF ratio + absolute gate → HF-band entropy confidence → soft-knee target gain → attack/release → split-band HF attenuation → optional Dry/Wet → optional output-level stage.

The same selected frequency range is extracted with Filter (pass Hann band) for processing. For the normal output:

output = input − HF × (1 − gain)

At gain = 1 the equation reduces exactly to the input. The low-frequency residual is not multiplied by the de-essing gain, so a sibilant event does not intentionally duck the fundamental, low formants or other content outside the selected HF band.

No random process is used; identical input and settings produce deterministic processing.

Channels & timing

Every input channel is processed and retained. By default, detection examines all channels and, frame by frame, uses the channel with the strongest valid HF evidence to drive one shared gain curve. Applying the same gain trajectory to every channel preserves inter-channel level relationships while still allowing a sibilant event on any channel to trigger the processor.

The Sound is processed on a temporary copy shifted to time 0, and the original start time is restored to the result. Sample rate, duration and number of channels are preserved.

Detector timing: the Intensity analysis uses a 200 Hz pitch floor. In Praat this corresponds to an effective intensity window of about 16 ms and, with automatic time step, approximately 4 ms between frames. Attack and Release therefore smooth a detector that already has this finite temporal resolution.

The script uses a conservative short-file guard and requires at least about 48 ms of input before running the detector.

Output & monitoring

Normal processed output

With Listen_to_removed = Off, the output is named:

<original>_deessed

Removed-only monitoring

With Listen_to_removed = On, the output contains only what the gain curve removed from the selected HF band:

removed = HF × (1 − gain)
name = <original>_sibilants

Dry/Wet is not applied to this monitoring output.

Output level modes

ModeBehavior
Natural levelNo final level scaling.
Match input RMSScales the complete output to the input Sound's RMS.
Safety ceilingAttenuates only when the output peak exceeds Ceiling_peak.
Peak normalizeAlways scales a non-silent output so its peak equals Ceiling_peak.

If the stored result peaks above 1.0, the script warns about possible clipping in integer PCM. For playback only, it creates a temporary copy scaled to 0.95; the stored Sound itself is not changed by that playback safeguard.

Visualization

The optional 8 × 8 Picture page contains:

The waveform display deliberately uses a real channel rather than a mono fold, so anti-phase multichannel material cannot disappear from the diagnostic plot through cancellation.

Further Reading