Why your denoiser swallows word endings, and how to measure it
Aggressive noise reduction files the tails off your words. Here is why classic denoisers do it, how to measure the loss instead of guessing, and how to clean without the damage.
You run a noisy recording through a denoiser and the hiss disappears. Then you listen closely and the words have lost their tails. Final consonants soften, breaths vanish, and the ends of words get swallowed into the silence the denoiser worked so hard to create. The noise is gone, and so is a little of the speech.
Why denoisers eat word endings
Most classic noise reduction is a gate on the frequency spectrum. It learns what the noise looks like when nobody is talking, then subtracts that profile everywhere. The trouble is that the quiet end of a word, an unvoiced s, f, or t, or the soft decay of a vowel, looks a lot like noise. It is low in energy and broad in frequency, so the gate treats it as something to remove. The louder middle of the word survives; the fragile edges do not.
How to measure the loss, not guess at it
“It sounds a bit off” is not something you can act on. Turn it into a number:
- Transcribe both versions. Run speech to text on the original and the cleaned file, then compare. A rising word error rate after denoising, especially on short function words, is word-ending loss showing up as dropped or mangled words.
- Watch the sibilants. Line up the two waveforms and look at the high-frequency energy on s and f sounds. If it is markedly lower after cleaning, the denoiser is cutting speech, not just noise.
- Keep an A and B. Always audition the cleaned file against the original on the same phrase. If a word feels clipped, it usually is.
Cleaning without the collateral damage
The fix is to be gentler and smarter. A lighter reduction that leaves a little noise but keeps the consonants is often better than an aggressive one that files the words down. Newer approaches resynthesize speech rather than subtract a spectrum, so they reconstruct the quiet edges instead of gating them away.
This is why Vocast treats word-ending preservation as a measured property, not a hope. Its cleanup is tuned to keep speech, including endings and breaths, and the result is checked rather than assumed. If you clean audio a lot, measure the loss first; you cannot improve what you have not looked at.