Guides

    Speaker IDs and Non-Speech in Captions

    A style guide for speaker labels, bracketed sound, italics for off-screen speech, and music cues that auto-captions drop and d/Deaf viewers actually need.

    Versely Team8 min read

    Auto-captions transcribe words. They drop the rest of the soundtrack: who is speaking, that the line is off-screen, that the room laughed, that the music changed key under the CTA. WCAG's definition of captions includes those equivalents. 47 CFR 79.1 calls the same thing completeness: speaker identification and non-speech information. d/Deaf viewers are not missing vocabulary. They are missing the layer ASR was never asked to emit.

    This is a style guide for putting that layer back. It follows the DCMP Captioning Key where the Key is specific, and it assumes you already have a speech transcript to annotate.

    Person editing caption text on a laptop

    Speaker IDs: name, then stop repeating

    The DCMP speaker-identification guidance is placement first, labels second.

    • When two people are on screen and the player can position a caption, put the line under the person who is talking. Placement is the ID. Do not also stamp (MAYA) on that line.
    • When placement cannot do the job (burned-in centered captions, a single lower-third style, a wide two-shot where both mouths move), put the name in parentheses on its own line above the dialogue:
    (Maya)
    I'm not refilling that.
    
    • Use the name the program uses. If the program has not named them, use a short role: (interviewer), (driver), (narrator). Do not invent "Woman 2" if (colleague) is clearer. Stay consistent for the whole file.
    • Re-ID after a gap, a new scene, or a third voice. Do not re-ID every line in a two-hander once the pattern is established.
    • Off-screen speech: if you know the side, DCMP allows placing the caption toward that side. If you cannot place, label and italicize the line.
    • Narrators get (narrator) once, then italics on the narration if italics are available, so narration does not look like on-camera dialogue.

    For generated UGC with one talking head, you often need zero IDs until a second voice (a VO, a second avatar, a phone call) enters. Add the ID on the first line of the new voice, not before.

    What auto-captioning does instead is concatenate both voices into one stream. Overlapping speech is the usual failure; timed captions from speech already flags overlap as a timing problem. It is also an identity problem. If two people talk at once, split the lines, label both, and do not let the engine interleave words into a single sentence.

    Brackets for non-speech, italics for off-screen

    DCMP's sound-effect rules are small and strict enough to memorize:

    • Caption a sound when it is necessary to understand or enjoy the media. Skip the fridge hum that is always there. Keep the glass that shatters, the phone that rings, the laugh the joke depends on.
    • Put the description in square brackets, lowercase: [glass shatters], [phone rings], [laughter].
    • Include the source in the brackets unless the source is clearly on screen. [door slams] when we do not see the door; [slams] is not enough. If we see the door, [door slams] is still fine; [slams] only works if the picture already owns the source.
    • Off-screen sound effects are italicized when italics exist: *[knocking]* or an italic [knocking], depending on the format. In WebVTT that is <i>[knocking]</i>. In a burn-in preset with no italic face, keep the brackets and skip the italic.
    • Never past-tense a sound. [dog barked] is a recap. [dog barking] or [dog barks] matches the moment.

    A short vocabulary that is enough for brand video:

    Sound Caption Skip when
    Laugh, from a person we see [laughter] The laugh is also in the dialogue line you already captioned as "ha"
    Laugh, off-screen audience <i>[laughter]</i>
    Music sting that punctuates a cut [music sting] Wall-to-wall bed (see music, below)
    Notification, off-screen <i>[phone chimes]</i> Decorative UI clicks
    Silence that is the joke [silence] Ordinary pauses between sentences

    Do not caption [background noise]. If you cannot name the source and the reason, it is not necessary.

    Music: identify it, do not dump the lyric sheet

    Music is non-speech. Lyrics, when they are plot or offer, are speech. DCMP splits them.

    • If a song is performing the soundtrack job (mood, sting, bed) and the lyrics are not the message, caption the music, not the words: [upbeat pop] or, when you have it, [upbeat pop: "Refill the Morning"]. Off-screen or background music is italicized.
    • If the lyrics are the message (a jingle that states the offer, a licensed chorus the sketch is built on), caption the lyrics as speech, marked as lyrics. DCMP uses music-note wrappers around lyric lines when the format can carry them. In WebVTT, a common pattern is a note in the first cue () and verbatim lines after. Do not dump six repeating chorus passes; caption the first and mark the rest [chorus continues] unless new words appear.
    • Instrumental beds under VO: one music cue when the bed starts, not a running commentary. [soft piano] at 0:00 is enough. Re-cue only if the music changes function (sting, drop to silence, a new track).
    • Do not editorialize. [sad music] is a judgment. [low strings] is a description. If the only honest caption is [music], use that rather than inventing a mood the picture did not earn.

    Auto-captioners either ignore music or hallucinate lyrics from a bed. Both are wrong. The human pass is: listen with the transcript muted in your head, and write the cues the transcript does not have.

    The pass after auto-caption

    Versely's auto-captions run speech recognition, chunk the words, and burn a preset. That is the right machine job. Speaker IDs and non-speech are the human job on top, using the same readability rules you already apply to line length and contrast.

    A working order:

    1. Generate the speech captions. Add captions to the video or ask the agent to transcribe and caption for the burn-in preview. In parallel, pull a plain transcript if you need a sidecar master. Speech-to-text is that step.
    2. Watch once with sound off and the captions on. Every time you cannot tell who spoke, mark an ID. Every time a sound does work the picture does not, mark a cue.
    3. Watch once with sound on and captions off. Confirm you did not caption decorative noise, and that music cues are not lyric dumps.
    4. Edit the sidecar (VTT) or the burned-in text. IDs on their own line. Brackets lowercase. Italics only for off-screen, narration, and off-screen sound/music.
    5. Check reading time. (Maya) plus [laughter] plus a full sentence in a 0.8-second cue will not be read. Split, or hold the cue. Completeness does not override readability.
    6. For a dub, rebuild this layer in the dubbed language; do not translate [laughter] cues off an English file and hope the laugh still sits there. Keeping captions in sync with a dub is the timing half of that job.

    Burn-in presets will not italicize on command the way WebVTT will. If the destination is a feed, you may only have brackets and speaker names in roman. That is still SDH content, carried as open captions. If the destination is a site or a TV ingest, put the italics and IDs in the VTT or the 608/708 master, not in a screenshot of the burn-in.

    A one-page sheet to tape next to the timeline:

    IDs
    - Place under speaker if you can.
    - Else (Name) on its own line.
    - Re-ID after a scene change or a new voice.
    - Off-screen: italics + ID.
    
    Non-speech
    - [source + sound] in lowercase brackets.
    - Italicize off-screen sounds.
    - Present tense. No [dog barked].
    
    Music
    - [description] when it starts or changes function.
    - Lyrics only when they carry meaning.
    - No mood adjectives you cannot hear as instrumentation.
    
    Do not
    - ID every line of a two-hander.
    - Caption the fridge.
    - Paste the full chorus four times.
    - Trust ASR on overlapping speech.
    

    FAQ

    Should I caption [music] on a track that never stops?

    Once, when it becomes relevant, and again if it changes. A continuous bed does not need a cue on every caption block. If the bed is the only audio (no VO), a music description plus any plot-pertinent lyrics is the caption track.

    What if I do not know the speaker's name?

    Use a stable role or a visible trait the file already uses: (host), (customer), (voice on speaker). Do not switch from (woman) to (Maya) mid-file unless the program names her at that moment; then ID the naming beat and use Maya after.

    Do burned-in social captions need this, or only sidecar files?

    Both, when the destination is doing accessibility work. A muted feed still has viewers who cannot hear, and they still need to know it was a laugh, not a gasp. Social burn-in has less room, so be harsher about what is necessary: IDs on the first line of a new voice, one music cue, sound effects that are the joke. Sidecar SDH for long-form can afford more.

    Will the auto-captioner add [laughter] if I ask nicely in the prompt?

    Not as a reliable behavior. Speech recognition is optimized for words. Treat speaker IDs, brackets, and music as an edit pass on the transcript, the same way you already fix product names and prices. The engine gets you the speech; completeness is still a person.