Set AD Level Against the Dialogue
AD that ducks the whole mix is unusable; AD under the music is inaudible. Match description to dialogue, keep HI and VI mixes apart, and check on a phone speaker.
The usual AD mistake on a brand film is not a missing script. It is a level. One version ducks the entire programme every time the describer speaks, so the room tone, the music cue, and the tail of the last line all collapse. The other version treats description as a shy commentary bed and parks it under the score. A listener who cannot see the picture gets either a pumping, disorienting mix or a voice they cannot catch. Both tracks "exist." Neither is usable.
How to write the description, where to place it in the gaps, and how to generate the voice are covered in audio description for brand video. This is the mix pass that comes after that script exists.
Two failures, one fader habit
Failure A: the global duck. A compressor or a dropped master rides the whole mix down 8 or 10 dB for every AD line. Dialogue that was still decaying gets swallowed. Ambience that told the listener they were still in the warehouse disappears. Music that was doing scene-change work flattens. The describer is loud, and the film around them is gone. This is often a sidechain with too much range, or a static original_volume pulled down so the new track "wins."
Failure B: AD as music. The describer is mixed like a bed: polite, low, "out of the way." On headphones you can follow it if you already know the pictures. On a phone speaker it is gone. The listener who needed the track is the person least likely to be on studio monitors.
WCAG 2.2 Success Criterion 1.2.5 (Level AA) requires audio description for prerecorded video in synchronized media when the existing audio does not already convey the visuals. The technique is spoken description, in the same language, fitted into pauses in dialogue. The criterion does not publish a dB offset. It does assume the description can actually be heard, and that it has not replaced the programme it is describing.
Netflix's published Audio Description Style Guide v2.5, in the technical requirements, is more specific because it has to be mixed against a printmaster:
- Dip the original-version mix 6–12 dB during descriptive events, at the mixer's discretion.
- AD/VO should be clear and intelligible, with the natural presence of original dialogue still underneath when they overlap.
- Transitions in and out of those dips should take no more than 5 seconds. No abrupt level jumps.
- Do not raise the AD voice above the loudness spec to beat a loud event in the printmaster. Dip the printmaster instead.
- On 5.1, dip the centre channel for descriptive events; left and right only when necessary, generally no more than 6 dB, sparingly up to 12 dB.
That is a Netflix partner spec, not a Versely control and not a WCAG number. Steal the shape: description sits at speech level; the programme under it is dipped, not erased; the dip is local to the event, not a new quieter film.
EBU R 128 still applies to the finished AD version as a programme: target −23.0 LUFS, true peak −1 dBTP. Adding a speech-level describer will raise integrated loudness if you do not make room. Make room in the music and effects during the AD line, not by crushing the whole timeline.
Target the dialogue, not the bed
Set the describer against a representative line of spoken dialogue, not against the music, and not against the programme's integrated loudness.
- Pick a dialogue anchor. A line in the middle of the film, not a whispered aside and not a shouted hook. Measure that line's short-term loudness, or match it by ear on the same monitors you will mix on.
- Match AD speech to that anchor. The describer should sound like another speaker in the same film, slightly more neutral, not like a podcast dumped on top and not like a whisper. Delivery is still the flat, unhurried read from the earlier AD piece. Level is dialogue level.
- Duck what the describer is competing with. That is music and effects in the gap, the same way you duck music under a voiceover. It is not the dialogue, because if AD is sitting on a line of dialogue you wrote the description in the wrong place. Netflix allows overlap as a last resort; WCAG's default technique is the pause. Brand films should prefer the pause.
- Hold the dip for the line, then release. A five-second-or-faster ramp, in Netflix's language. A hard mute of the score at the first syllable is Failure A with extra steps.
- Do not "fix" a loud action cue by making AD louder than dialogue. Dip the cue. Loudness normalisation is why a quiet AD version will also play quieter than everything around it on a platform that only turns hot files down.
If the music still masks the describer after a sensible dip, the bed is fighting for the same frequencies as speech. Split it. Stem separation and splitting a song into stems get you an instrumental you can actually duck. Pulling a vocal hook out of the bed is often the whole fix.
HI mix and VI mix are different deliverables
Broadcast has treated these as two different mixes for a long time. They serve different ears. Brand teams collapse them because both get filed under "accessibility audio."
| Mix | Who it is for | What it does | What it must not do |
|---|---|---|---|
| HI (hearing-impaired / speech-forward) | Listeners who can see the picture but struggle to catch words | Dialogue up, music and effects down, sometimes mild speech enhancement | Add a describer. The viewer can see the pictures. Extra talk is noise. |
| VI (visually-impaired / AD) | Listeners who cannot rely on the picture | Original (or near-original) programme, plus description at dialogue level in the gaps | Strip the score and effects. Those are how the listener stays in the scene. |
An HI mix labelled as AD is a trap for both groups. A hard-of-hearing viewer gets a describer talking over residual dialogue they were trying to follow. A blind viewer gets a thin, over-ducked soundtrack and a describer that may still be too quiet because someone mixed it "politely" on top of an already-reduced bed.
Ship them as separate files when you ship both. The AD version is the VI mix: programme intact, description in the pauses, music dipped under the describer only. The HI version, if you make one, is a speech-forward programme without AD. Do not average them into one "accessible audio" export.
WCAG 1.2.5 is the AD/VI requirement. It is not a mandate to destroy the music. Extended description (SC 1.2.7, Level AAA) pauses the picture to make room when there are no gaps. That is a timeline decision, not a fader decision. Do not fake extended AD by talking over dialogue at matching level.
Mix it in the tools you actually have, then check a phone
Versely's attach_audio_to_video mix is static. music_volume and original_volume hold for the whole clip. There is no sidechain that ducks only while the describer talks. That is the same constraint as mixing a bed under a voiceover, and it changes the procedure:
- If the AD lines are sparse and the bed is already low enough for dialogue, attach the describer in mix mode with
original_volumeleft at the dialogue setting and the AD file itself riding at that same speech level. You are adding a voice, not turning the film down. - If the bed masks the describer in the gaps, lower the bed (or swap in the instrumental stem) for the AD export. That is a separate mix, which you wanted anyway. Do not lower
original_volumeto make AD audible: that ducks the dialogue and the ambience together, which is Failure A. - If you need a true dip only in the AD gaps, you cannot get that from one static pair of sliders. Build the AD version as its own mix: programme with the bed already reduced, then attach AD at dialogue level. Or mix the AD file against a ducked programme in a tool that has automation, then lay that stem back. The generate-and-attach path is add a voiceover to a video; the voice itself comes from voice-over. Pick a neutral read, not the expressive one you used for the hero VO. The best text-to-speech model ranking is a starting point for that choice.
Then do the phone-speaker check. Not Bluetooth earbuds. The built-in speaker, at a volume you would actually use at a desk.
- Play a dialogue scene. Confirm you can follow the words.
- Play an AD gap on the same speaker, same volume. Confirm you can follow the describer and still hear that a scene is happening underneath.
- If the describer vanishes, it is Failure B. Raise AD toward the dialogue anchor, or dip the bed, or both.
- If the scene underneath vanishes, it is Failure A. You ducked too much. Put programme back until the warehouse, the street, or the office is still a place.
- Repeat once on the laptop speaker. Phone transducers hide quiet speech first; a pass on both is the point.
If you cannot hear the difference between HI-style speech-forward and VI-style AD on that speaker, you have not made two mixes. You have made one compromised file and named it twice.
FAQ
Should AD be quieter than the dialogue so it does not "compete"?
No. Competition is a writing problem (you are on top of a line) or a bed problem (the score is too loud in the gap). The describer is speech. Match it to dialogue. Netflix's partner guide wants AD clear and intelligible, with original dialogue still present when they must overlap. Polite-and-quiet is how AD dies on a phone.
Is a 6–12 dB dip a Versely setting?
No. That range is from Netflix's Audio Description Style Guide v2.5, for dipping a printmaster under description. Use it as a reference for how much the programme should yield during an AD line, not as a number you type into music_volume. Versely's attach step is a static relative level for the whole clip.
Can I ship one mix that is both HI and AD?
You can put both treatments in one file. You should not. HI listeners asked for clearer dialogue, not a narrator. VI listeners asked for description without losing the scene. One file that ducks the programme and adds a quiet describer fails both. Two exports, labelled, is the broadcast pattern and the one that still works on a brand site.
What if there are no pauses for AD?
Then a simple mix will talk over dialogue, which WCAG's default technique does not want and which Netflix treats as a last resort. Either recut the film to create gaps, or use extended description (pause the picture, describe, resume), which is SC 1.2.7 at Level AAA. Turning the whole mix down so two people can talk at once is not a substitute.