Check Dialogue Intelligibility Before Sign-Off
Pass a speech-to-noise check and a mono phone test before sign-off, and fix masked words with EQ and reverb moves rather than raising the VO.
A mix can hit its loudness number and still lose the line in a car or a kitchen. Integrated LUFS does not know whether the words survived. It knows whether the programme was loud. A bed that eats consonants, a reverb tail that smears them, and a stereo image that collapses on a phone speaker will all meter legally. Sign-off that only looks at a loudness meter is how those mixes ship.
The check that catches this is not "turn the voice up." It is a speech-to-noise listen against a real noise floor, a cheap mono-phone test, and then EQ and reverb moves on the bed and the room, not another 2 dB on the VO fader. Raising the voice after the mix is already loud just spends the loudness budget on the same masked words.
Meters pass mixes that kitchens fail
Broadcast and podcast sheets measure programme loudness and true peak. They do not measure whether a sentence is intelligible with a kettle on. IEC 60268-16 is the standard that does define an objective speech-intelligibility rating, the Speech Transmission Index (STI), for rooms and sound systems. A mix review is not an STI survey: you do not have the test signal, the source position, or the three-measurement average the standard describes. What you can steal from that world is the idea that intelligibility is a separate pass from level.
The practical version, for a generated VO under a generated bed, is a speech-to-noise check:
- Solo the voice. Confirm the words are clear in isolation. If they are not, stop. Isolation, a re-read, or a drier generate comes first. Rescuing a noisy voiceover and keeping reverb out of generated dialogue are the two usual upstream fixes. A wet native-audio take will not become intelligible because you ducked the music.
- Put the bed back at the mix level. Do not solo-then-guess; listen to both.
- Add a kitchen-or-car noise floor. A recording of a cooker hood, a cabin at 50 mph, or just the room you are sitting in with a tap running is enough. Playback the mix at a modest volume over that floor, not in a treated booth.
- Write down every word you cannot catch without leaning in. Those are the fails. "I know what it says because I wrote it" is not a pass.
If the voice is a clean TTS bounce and the bed is a finished music generate, the fail is almost always masking, not a bad read. A voiceover that sounds muddy under AI music is the 250–500 Hz collision. This check is how you confirm you actually fixed it, including the words that only disappear in a car.
Do not use captions as a substitute. Captions are a parallel path for the same information; they do not prove the soundtrack is intelligible. If the piece will be watched with sound on, the line has to work with sound on.
The cheap mono-phone test
After the noisy-room listen, do the device listen. Most of the audience is one small speaker, in mono.
- Bounce a stereo mix. Sum it to mono, or use the phone's speaker, which will sum it for you.
- Play it on a phone at arm's length, 30–50% volume, in the same ordinary room. Not on the studio monitors, not on AirPods.
- Confirm the line, not the bed. If you can hum the music and cannot repeat the last sentence, the fold or the mask failed.
- If the line dips only in mono, the problem is width or phase, not level. Collapse the bed (mid/side, pull the side down) and retest. Do not reach for the VO fader to fix a cancellation.
This is the same physical check used for spatial audio, pointed at words. Headphones will lie: a wide, wet, loud mix feels expensive and still dies on a speaker. The editor's 480p preview pass is free, with a short per-user cooldown, and audio survives the downscale, so you can run this on a preview. The final export is charged once. A preview on a phone speaker is a valid intelligibility pass. A preview on studio monitors is not.
Have someone who did not write the script do step 3. You will hear through the holes because you know the copy.
Fix the words without raising the VO
If the check fails, the instinct is to ride the voice up. That is the last move, not the first, because the programme is probably already near its loudness ceiling. Turning the VO up 3 dB either pushes true peak over the cap or forces a limiter that then squashes the consonants you were trying to save. The bed comes up with the programme when you "make it louder." The mask remains.
Work this list in order. Stop at the first pass.
1. Pull the bed down, as a relationship, not a vibe. On Versely, attach_audio_to_video in mix mode is a static music_volume versus original_volume for the whole clip. Ducking the music under a voiceover is that control. Pull further than feels fair when the bed is soloed. A working range from the muddy-VO piece is the bed 18–25 dB under the voice. If the bed disappears in the gaps and you need it to bloom, that is an NLE sidechain after this pass, not a reason to leave it loud under speech.
2. Carve the bed, not the voice. Cut 250–500 Hz on the music, wide, 3–6 dB. High-pass the bed around 80–100 Hz if you do not need musical bass. Tame 4–7 kHz on the bed if generated hats are fighting sibilance. Leave the VO's body alone. Carving the voice to make room for a dense generate is how the line gets thin and then you turn it up, which is the failure mode this list exists to avoid.
3. Kill the competing vocal. If the generate has a sung hook, split it. separate_music_vocals on an in-app generate; isolate on an uploaded mix. Use the instrumental as the bed. Two voices is not an intelligibility problem you can EQ out.
4. Dry the voice before you EQ it. Reverb on generated dialogue is baked in, not a send. If the take is wet, a new dry read (or a prompt that names a close, dry mic) beats a de-verb that chews consonants. Presence EQ cannot restore a smear. If the voice is already dry and still dull on a phone, a small presence lift on the VO around 2–4 kHz is legal. A high shelf on the whole mix is not; it will hype the bed's grit too.
5. Centre the line. A pan or a spatial native-audio move that puts the words off to one side will lose them in the phone sum. Recentre. Width belongs to the bed, and only as much width as the mono test allows.
6. Only then raise the VO, and only if the programme still has loudness headroom. Measure integrated after the move. If you were already at the destination's target, raising the voice means lowering the bed further or accepting a hotter programme. The first is usually right.
Adding the voiceover in mix mode, not replace, is what keeps a useful original (room, native audio, production effects) under the line. Replace is for when that original is the thing masking the words. The publish check should include "phone speaker, mono, kitchen noise" as a tick, not only captions and duration.
A mix that fails this check and still ships because the LUFS number was right is not a loudness success. It is an intelligibility skip. Put the phone test on the sign-off sheet next to the meter screenshot so they are the same conversation.
FAQ
How much louder should the voice be than the bed?
Loud enough that a stranger can repeat the line on a phone speaker in a noisy room. The 18–25 dB-under working range is a starting relationship for a static bed, not a published standard. If the words fail the phone test at that relationship, the bed is the wrong shape (carve, stem, or a sparser generate), not merely the wrong fader.
Can I skip the phone test if I mix on cheap speakers?
Cheap monitors in a quiet room are not a kitchen, and they are not a 3-inch mono speaker at arm's length. Use the phone. It is the device. A mix that only works on NS-10s has the same problem in reverse.
Does turning on captions fix intelligibility?
Captions fix access when the sound is off or the viewer needs text. They do not fix a soundtrack. If the brief is sound-on, the line has to work sound-on. Run both: captions for the mute path, phone test for the speaker path.
When do I re-generate instead of mixing?
When the voice itself is wet, noisy, or the bed is a dense mastered record with a sung hook you cannot stem. Isolation and a carve have a ceiling. A dry TTS bounce and a sparse instrumental generate are often cheaper, in time, than another hour of EQ on the wrong files. Re-generate the source, then run this check again. Do not sign off a "good enough" mask because the meter went green.