Write Audio Description That Fits the Gaps
Present tense, character IDs, and dialogue-gap timing for audio description, plus a timed script template you can fill for a 30-second spot.
Standard audio description is not a second voiceover. It is lines that fit inside silences the picture already has. WCAG 2.2 Success Criterion 1.2.5 is explicit: during existing pauses in dialogue, the track describes actions, characters, scene changes, and on-screen text the main soundtrack does not already convey. If the line does not fit the gap, it is not standard AD. It is a collision. The case that AD is missing from brand video is already made; this post is the script.
Present tense, present continuous, no reviewer's voice
Netflix's Audio Description Style Guide v2.5 wants description that is informative, conversational, present tense, and third-person omniscient. The DCMP tip sheet matches: present tense, active voice, third person. Past tense recaps. Future tense spoils.
Use simple present for a beat that lands in a gap ("She holds up the bottle"). Use present continuous when the action continues across the gap ("Rain sheets across the windshield"). Do not stack both on the same beat.
Write what is visible, not what you inferred:
| Avoid | Write |
|---|---|
| She looks sad about the price. | Her mouth tightens. She lowers the bottle. |
| A stylish millennial in a loft. | A woman in her thirties, dark coiled hair, stands at a white worktop. |
| The product is premium. | Brushed aluminium, a matte black pump. |
| He angrily slams the laptop. | He shuts the laptop. |
Netflix's verb rule is the compression tool: a specific verb instead of a bland verb plus adverb. "He hobbles" is shorter than "he walks with difficulty."
Do not narrate camera grammar unless the grammar is the information ("a close-up of the ingredient list"). "We smash cut to" is a reviewer's note. Directional language is from the viewer's body ("to the right of the sink"), not production slang.
Leave room for dialogue, music, sound effects, and intentional silence. Plot-pertinent speech wins the gap. Talking over the hero line is an extended-description problem, not a writing-faster problem.
Character IDs that do not leak the script
Identify people the way the picture does, in that order.
- Appearance before the name, unless the soundtrack already named them. Netflix default: unnamed until dialogue or a plot point introduces them.
- If you must name early (a crowd, a timing bind), descriptor first: "a bearded man, Jack," not "Jack, who has a beard."
- Do not name a figure the film is withholding.
- Once named, stay consistent. If she is "Maya" at 0:04, she is not "the woman" at 0:18 unless a group has swallowed her.
- Pronouns only with a clear antecedent. Two people in frame plus "she turns" is a guess.
- Main and relevant supporting characters. Skip extras unless the extra is the shot.
- Visible traits (hair, skin, build, age band, visible disability) when they identify the person, applied consistently. Netflix asks for person-first phrasing ("a swimmer with one leg") and forbids guessing race, ethnicity, or gender the picture and plot have not established.
On a 30-second spot you get one chance to plant an ID. Spend it on the person whose action carries the offer, then stop re-describing clothes.
Fitting the line into the gap
The gap is a measured silence. Build against a transcript and a picture log, not memory.
- Pull a plain transcript of the soundtrack. Mark every stretch with no speech. Keep a breath if it is long enough to use; drop it if a sound effect owns the beat.
- Watch those stretches only. List visuals the audio does not already cover: who, what changed, on-screen text, product state, location shift.
- Rank. If the listener would miss the offer without it, it stays. If the VO already said it, it goes. Decorative b-roll goes.
- Write one line per gap. Say it out loud against the picture with the soundtrack up. If you overlap the next word of dialogue, cut the line. Do not speed-read.
- If a necessary fact has no gap, do not squeeze it into standard AD. Flag it for an extended cut, or move it into the main VO on the next picture revision.
If you cannot speak the line calmly inside the silence, it is too long. Cut filler ("now," "we see," "the camera") first. On-screen text that carries the offer is a fact, and it still has to fit.
Voicing is a separate pass. AD informs; it does not perform. In Versely, generate_speech takes style_instructions; use them for a neutral, unhurried read. Mix with attach_audio_to_video in mix mode so the original track stays. Voice-over and adding a voiceover to a video are that generate-then-mix shape, pointed at a second track. Pick a text-to-speech model that can deliver flat and clear; character-forward reads fight the format.
A timed template for a 30-second spot
Copy the table. Fill it from the finished picture cut, not the board. Times are mm:ss.s. "Gap" is the silence you may speak in. Treat it as a cap, not a target to fill.
Spot: 30-second refill ad, woman on camera, VO already in the soundtrack. Versely's default picture rate is 25 fps; keep AD timecodes on that clock so an editor does not relink against 24.
| In | Out | Gap | Picture (only what is not in the VO) | AD line | Words |
|---|---|---|---|---|---|
| 0:00.0 | 0:03.0 | 3.0s | Empty worktop. A hand sets down a tall matte-black pump bottle, label facing us. | A hand sets a black pump bottle down. | 8 |
| 0:03.0 | 0:08.0 | n/a | She speaks on camera. | (no AD) | 0 |
| 0:08.0 | 0:10.0 | 2.0s | Cut to sink. She rinses a cloudy glass bottle and stands it in a rack. | She racks the rinsed bottle. | 5 |
| 0:10.0 | 0:16.0 | n/a | VO continues over refill pour. | (no AD) | 0 |
| 0:16.0 | 0:18.4 | 2.4s | Close-up: a super, "30-day refill · $24", lower third. VO is between sentences. | On-screen: 30-day refill, 24 dollars. | 7 |
| 0:18.4 | 0:24.0 | n/a | VO. She smiles, pumps once. | (no AD; smile is already in her voice) | 0 |
| 0:24.0 | 0:26.5 | 2.5s | End card. Logo. URL in type only: refill.example.com. No speech. |
Logo. Refill dot example dot com. | 6 |
| 0:26.5 | 0:30.0 | 3.5s | Legal two-liner in 8-pt type. Music only. | One refill covers 30 days. Cancel anytime. | 8 |
In / Out come from silent ranges, not shot changes. A cut inside a VO sentence is not a gap. Picture is a writer note, never spoken. AD line is what you record, present tense, no "we see." Words is the check; if the count climbs, cut before you record.
The 0:16 super and the 0:24 URL are the lines teams skip and the listener needs. If either collides with VO, do not whisper faster. Move the super, or fold the words into the hero VO and drop them from AD.
Then: spoken column to a script, generate the read through the agent voiceover flow or generate_speech, mix onto a copy of the picture, listen once with eyes closed. If you lost the offer, the script is short. If you lost a sentence of the original VO, the script is long.
FAQ
Can I write AD in the same document as the hero VO?
Yes, as a two-column script (VO | AD) with shared timecodes. Do not let the VO writer fill AD as an afterthought in the VO column. Different job, different tense, different permission to speak.
How do I ID a mascot or a generated character that is never named on the soundtrack?
Describe the figure as the picture shows it, then reuse a short consistent label ("the fox in the red coat"). Do not invent a proper name the film did not earn. If the brand name is on screen and unspoken, that is on-screen text: read it in a gap.
What if the 30-second cut has no gaps at all?
Standard AD cannot be written into a wall of speech. Flag every necessary visual that has no silence, and either recut the picture to open gaps, fold those facts into the hero VO, or schedule an extended-description version that freezes the picture. Forcing lines over dialogue is how AD becomes a second ad talking on top of the first.
Does the AD voice need to be a different person from the hero VO?
Not necessarily a different identity, but a different delivery. Hero VO persuades. AD reports. Distinct style_instructions on the same text-to-speech model are usually enough on a 30-second spot.