Guides

    De-Ess Generated Voices Without a Lisp

    TTS and clones hiss in a way live VO does not. A split-band de-esser setup, and when the real fix is a different take or engine rather than more processing.

    Versely Team9 min read

    Generated S is a different problem from a live voiceover that got a little bright. Text-to-speech and many clones put a hard, consistent edge on S, SH, and CH: the same frequency, the same amount, every time the phoneme appears. A live reader sprays a bit, then backs off. A model does not get tired and does not hear the take. If you treat that edge with a wideband de-esser, the whole syllable ducks and the line turns into a lisp. Split the band. If the edge is still a weapon after a modest split-band cut, the fix is a different take or a different engine, not a third plugin.

    Versely does not de-ess. Generate the line, listen, then process in a DAW, or generate again. Isolation can make this worse: pulling a voice out of a mix often hyps the same presence band the model already favoured.

    Why generated S is sharper than a live take

    Three stacked causes, and they add.

    The model likes intelligibility. TTS is trained to be understood through small speakers. Presence and sibilance are cheap intelligibility. The result is a voice that cuts through a bed on headphones and whistles on a phone.

    The take is identical every time. A human's fourth read is softer on the esses. A model's fourth read of the same text is the same S, unless you change text, voice, or engine. Useful for matching a campaign. Brutal when the S is wrong.

    Clones copy the source's S and then add the engine. Voice cloning will reproduce a sibilant donor. Isolation or a "clarity" EQ on top stacks the same band. Changing the voice on an existing clip keeps timing and swaps timbre. It does not remove a lisp you already baked in with EQ.

    Reverb is a different defect. Room is convolved into the take and does not come out later. Sibilance is usually a band problem on an otherwise usable take. Try a split-band pass and a re-generate before you throw a dry read away. Do not keep a wet, sibilant read and hope two processors will make it a person.

    A quick listen to name the fault:

    What you hear Likely cause First move
    Every S is a short whistle, VO otherwise fine Narrow sibilant peak Split-band de-ess, sweep 4-10 kHz
    Whole syllable "thoftens" when you de-ess Wideband reduction Split the band; reduce range
    S harsh only on earbuds / phone, fine on monitors Translation, not a DAW fail Phone pass, then less range than you used on headphones
    S harsh and the voice is thin or fizzy all the time Wrong voice or engine, or a bright EQ already applied New take / new voice_id, do not stack de-essers
    S harsh only after isolation Isolation hyped presence De-ess after isolation, modestly; do not pre-brighten

    Split-band, not wideband

    A de-esser is a compressor with a trigger in the sibilant range. The damage depends on what it turns down.

    Wideband. The detector hears an S, the gain reduction applies to the whole voice. The body of the syllable drops with the hiss. That is the lisp: "sales" becomes "thales," not because you EQ'd in a lisp frequency, but because you punched a hole in the whole phoneme.

    Split-band (or "split," "multiband," "ess-only"). The detector still hears the S. Gain reduction applies only to the band that contains it. The vowel around it stays put. That is the processor you want on generated speech.

    If the plugin has a listen key for the detector, use it. Sweep until the whistle is all you hear, then switch back to reduction. Generated S is often a narrow peak: a narrow split-band cut takes the whistle and leaves the voice. A wide 3 kHz-to-12 kHz "de-ess" is a presence cut in costume.

    Do this on DX only. Never de-ess the printmaster. You will duck hats and UI ticks that coincide with an S, and think the voice got worse.

    Do it after pitch/time, isolation, and any "clarity" EQ. De-essing first and then boosting 6 kHz puts the whistle back. De-essing the printmaster after a limiter can make the detector chatter on the limiter's own crunch.

    A starting setup you can hear

    These are starting points. Sweep, then listen on a phone speaker. Headphones hide nothing and also lie about how bad the S is.

    1. Insert a split-band de-esser on DX. If you only have a wideband one, use an EQ with a dynamic band instead: a dynamic cut whose detector is the same band it cuts.
    2. Find the whistle. Solo the detection band. Sweep between about 4 kHz and 10 kHz while the line plays. Generated voices often sit in a smaller window inside that range. The correct frequency is the one where the S is a spike and the rest of the words almost disappear from the solo.
    3. Set range, not smash. Start around 3-6 dB of reduction on the peaks. If you need more than that, the take is the problem.
    4. Attack fast enough to catch the burst, release fast enough that the next vowel is not dull. If the release is slow, every S leaves a dull hole behind it. That hole is another lisp.
    5. Leave a little S. A voice with no sibilance is a lisp in the other direction, or a wet towel. Toggle the processor. You should hear the whistle go, not the person go.
    6. Phone-speaker check. If the S is gone on monitors and screaming on the phone, you under-did it. If the phone sounds lispy, you over-did it. Trust the phone more than the headphones for this defect.

    A static EQ cut at the whistle can work when every S is the same, which is how TTS behaves. A dynamic band is kinder when some esses are worse than others. Do not stack a static cut, a de-esser, and a "podcast" preset.

    Script is a processor too. Lines packed with s-clusters ("success starts Saturday") are a gift to a sibilant engine. Rewrite before you add a second de-esser. Delivery tags are not de-essers: [whispers] changes delivery. It does not notch 7 kHz.

    When processing is the wrong fix

    Stop adding processors when any of these is true.

    The reduction needed is huge. If a split-band pass at a sane range still leaves a whistle, or a larger range lisps, the voice is the wrong voice. Pick another voice_id, or another engine. Generate the same sentence twice and keep the one whose S you would ship without a plugin.

    The harshness is the whole timbre. A fizzy, always-bright clone will not become a warm reader because you notched the esses. Clone from a better sample (closer, drier, less already-processed), or use a catalog voice. A processed Instagram VO as a donor is how you clone a de-esser's leftovers.

    You already EQ'd "clarity" into DX to beat a muddy bed. That fight belongs on the bed (carve 250-500 Hz on MX, pull the bed down). Brightening DX to win a masking war creates a sibilance job. Un-do the presence boost, carve the MX, then see if you still need a de-esser.

    The take is wet. De-essing a reverberant S smears the room's own treble. Get a dry generate: close-mic language in the prompt if it is native audio, or TTS, which you can keep dry by not asking for a hall. Process S on a dry DX.

    You are de-essing a mix because there is no DX stem. Split first if you can (isolate_audio, or the original TTS file you should have kept). De-essing a printmaster ducks the bed on every S. That pumping is a new artefact, and you will try to fix it with more compression.

    Regenerate through the voiceover hub when the line is short. A 12-second VO is cheaper to do again than to "save" with three stacked processors. Change one variable per attempt: voice, then engine, then wording.

    Write and generate the voiceover when you are still choosing a voice. Listen to the S on a phone before you attach it under music. Once it is under a bright bed, every diagnosis gets harder.

    FAQ

    Why does de-essing make the voice sound like a lisp?

    The processor is turning down the whole syllable, or too much of the presence band. Switch to split-band, narrow detection until you only hear the whistle in solo, and cut the range until the S is a consonant again. If a modest split-band pass lisps, the take is the whistle. Change voice.

    Should I de-ess before or after adding the music bed?

    After you have a dry DX you like, de-ess that. Then mix the bed under it. If you mix first, you will de-ess hats by accident or miss the S because the pad masks it on monitors. Phone speakers will still play the S. DX first, then MX.

    Can I ask Versely to de-ess in the editor?

    No. Speech generate, voice change, and mix-mode attach do not include a de-esser. Generate a drier, darker, or different voice, or process DX in a DAW. A 480p picture preview (free, with a short per-user cooldown) will not tell you whether the S whistles on a phone.

    Is a darker voice always the fix?

    No. Some dark voices still spike on S, and a dark voice under a dark bed disappears. Intelligibility lives in the same neighbourhood as sibilance. The goal is a spike you can notch, not a muffled read. Choose a voice whose S you can live with at 3-6 dB of split-band reduction, on a phone, under the actual bed.