Guides

    A Hard-of-Hearing Mix as a Separate Deliverable

    An HI mix is not louder voiceover. Dialogue-forward EQ, reduced music, and a stem recipe for shipping HI and VI tracks instead of relying on captions alone.

    Versely Team9 min read

    Turning the voiceover up is not a hard-of-hearing mix. It is a louder master, which hits the limiter, which flattens the consonants you were trying to save. An HI mix is a different balance: dialogue in front, music out of the way, narrative effects kept, bed ambience dropped. It is also a different file. Shipping it as "the mix, but hotter" is how you fail both the HI viewer and the version everyone else was supposed to hear.

    Captions remain mandatory for prerecorded synchronized media under WCAG 2.2 SC 1.2.2. An HI mix does not replace them. People who cannot hear the dialogue still need the words on screen. People who can hear some of it, and who are fighting the music, need the soundtrack changed. Those are two audiences. One overlay does not serve both.

    HI is not VI, and neither is "louder"

    HI (hard of hearing) is a dialogue-forward soundtrack. The picture stays the same. The music and the beds move down so speech is the thing the ear can hold.

    VI (visually impaired) is audio description: a narration track that speaks the picture during existing pauses, aimed at people who cannot see the product, the on-screen price, or the logo. WCAG SC 1.2.5 is that job, at Level AA, and it is a different deliverable.

    Do not fold them. An HI mix with a shouted VO on top of a busy bed still hides consonants. A VI track mixed like an HI track still talks over the dialogue it is supposed to sit between. Buyers who ask for "an accessible mix" usually mean one of these and have not named it. Ask which.

    Captions are a third object. They do not make the soundtrack easier to hear. They make it readable. Ship all three when the brief is actually "this has to work for deaf, hard-of-hearing, and blind viewers." Ship HI + captions for most brand and training video. Add VI when the picture is carrying information the soundtrack never says.

    Dialogue-forward, not gain-forward

    The BBC's own mixing guidance is blunt about the music, not the fader on the voice. In its Best Practice Guide: Sound Mixing for BBC Programmes it reports that reducing music levels a little lets people across the audience, including people with certain types of hearing loss, hear dialogue better. After the full mix is complete, the BBC recommends taking the music down 4 dB. The same note warns about percussive sounds and lyrics under dialogue. That 4 dB is a starting cut on the music stem, not a 4 dB boost on the VO.

    Why not just raise the voice:

    • Programme loudness is measured on the mix. EBU R 128 targets −23.0 LUFS with a true-peak ceiling of −1 dBTP, on top of ITU-R BS.1770. A hotter VO with the bed left alone pushes integrated loudness up. A platform that normalises down will pull the whole file, music included, and you are back where you started with extra distortion. The existing write-up on why a video sounds quiet is this mechanism from the other side.
    • Hearing loss that makes speech hard in noise is often about masking in the presence band, roughly the 2–4 kHz region where consonants live. A bed with hats, synth arpeggios, or a sung hook sitting in that band will mask /s/, /t/, /f/ and /k/ even when the VO meter looks healthy. Cutting the bed, and cutting it in that band, does more than adding 2 dB to the voice.
    • A limiter on a boosted VO clips the transients that carry intelligibility. You hear "louder" and lose the letters.

    So the HI move is: leave the dialogue level near where the full mix had it, drop the music, thin the bed in the presence band, keep the effects that explain the picture (a click, a notification, a crowd reaction), drop the effects that are wallpaper.

    Static ducking, the kind already documented for voiceover beds, is the everyday version of this on a single file: attach_audio_to_video in mix mode with music_volume set meaningfully lower than original_volume for the whole clip. An HI deliverable is that idea taken further, on purpose, and exported as its own master rather than as the only master.

    When to ship HI / VI tracks versus relying on captions

    Viewer need Captions HI mix VI / AD track
    Cannot hear the dialogue at all Required Does not help Only if they also cannot see
    Can hear some speech, music masks it Helps as backup The actual fix Not the job
    Can hear, cannot see the picture Not the job Not the job Required when visuals carry meaning
    Sound-off autoplay on social Burn-in captions, different problem Pointless (there is no audio) Pointless

    Rely on captions alone when the destination is a muted feed, or when you do not control the soundtrack the player will use. Ship an HI mix when the destination is a page, a course player, a connected TV, or a buyer who asked for "clean audio" / "TV mix" / "hearing-impaired mix." Ship a VI track when the picture is doing work the VO never says out loud: a UI walkthrough, a product held up with a price, a logo sting with no voice.

    Broadcast and streaming players sometimes expose HI and VI as alternate audio. Most marketing pages do not. On the web, "a separate deliverable" usually means a second file (or a second <audio> track the player can switch), labelled in the surrounding copy, not a hidden stem inside one download.

    Stem recipe

    This is the pass. Do it on a copy of the timeline, not on the full mix.

    1. Get three stems, not one louder bounce. Dialogue. Music. Effects. If the soundtrack is a generated song, split it into vocal and instrumental stems (separate_music_vocals, tied to that generation's taskId and audioId). If the voice is buried in a found recording, isolate the vocal (isolate_audio on an audio_url). Those are two different tools. Do not feed a found interview into the generation splitter.

    2. High-pass the dialogue. Roll off rumble below about 80–100 Hz so the voice is not carrying energy that fights the limiter and does not carry consonants. This is mixing hygiene, not a published HI spec.

    3. Presence, not gain. A small lift on the dialogue in the 2–4 kHz region, or a small cut on the music in the same region, or both. You are unmasking letters, not making the file hotter.

    4. Music down 4 dB from the full-mix level, as the BBC starting point. If the bed still has a sung hook or a hat pattern under speech, drop it further or swap to the instrumental stem so you are not fighting another voice. Lyrics under dialogue are called out in the same BBC note for a reason.

    5. Effects: keep the ones that mean something. A UI click, a timer beep, an audience reaction that the captions will also mark. Drop room tone and "cinematic" whooshes that were only there for taste. HI listeners should not have to hear the grade.

    6. Re-measure the HI bounce. Same R 128 target as the full mix, same −1 dBTP ceiling. If the HI file is louder than the full mix, you raised the voice instead of cutting the bed. Bring it back.

    7. Export labelled files. project_fullmix.wav, project_HI.wav. If you also have description, project_VI.wav as a mix of the full (or HI) soundtrack plus the description VO sitting in the gaps, which is the generate-then-mix path on voice-over and replace video audio for the described version. Do not overwrite the full mix.

    8. Captions on every version. The HI bounce still needs a caption file or a burn-in. Intelligible audio is not a transcript.

    Iterate the HI bounce on the editor's free 480p preview pass (preview: true), which carries a short per-user cooldown, and spend the export charge once on the file you will actually hand over.

    A one-line agent request that matches the tools:

    Isolate the voice from this soundtrack. Mix it back over the instrumental stem with the music well under the voice, and give me a separate export. Do not replace the original mix.

    If there is no instrumental stem and the bed is original music you generated, split first, then mix. If the bed is someone else's commercial track, you have a rights problem, not a mixing problem, and isolation does not clear it.

    FAQ

    Can I ship one file with HI as the default?

    Only if the buyer asked for HI as the primary soundtrack. A course player or a public-service film sometimes wants that. A brand spot usually wants the full mix as default and HI as the alternate, because the full mix is the creative. Defaulting everyone to HI quietly throws away the music the spot paid for.

    Does a 4 dB music cut always fix masking?

    No. It is a published starting point, not a guarantee. A lyric hook or a hi-hat line in the presence band can still mask speech at −4 dB. If consonants are still gone, cut more, EQ the bed, or remove the competing vocal stem. Re-listen on small speakers, not on the studio volume you mixed at.

    Is an HI mix a substitute for captions on social?

    No. Social playback is often muted. An HI mix never plays. Burn-in captions are the social accessibility tool. HI is for destinations where the sound is actually on.

    Do I need a different voice for the HI mix?

    No. Same performance, different balance. A different voice is a recast, and it creates two "canonical" reads. Keep the take. Change the bed.