Workflows

    Music stems before you duck the voice

    A full mix under VO is a fight. Get stems, or generate the bed without a lead, then duck.

    Versely Team6 min read

    Ducking a full mix under a voiceover is asking a sidechain to solve an arrangement problem. The singer is still in the file. The kick is still in the same band as the chest of the voice. The limiter on the bounce is still holding the track up as if it were the record. Pull the bed down 12dB and the hook reappears every time the VO breathes. That is not a duck setting. That is the wrong file.

    Get stems first — or generate a bed that never had a lead — and then duck. This is a generation-side decision. It is not a loudness-target post and it is not a true-peak walkthrough. Those numbers belong on the last bounce, after the parts are right.

    ElevenLabs Music stems for an edit is the tool map. This is the order of operations.

    The fight, in one sentence

    A stereo song is mixed to be complete. A VO needs a hole. Completeness and a hole are opposites. Ducking changes level over time; it does not remove a part. If the part that collides with speech is still in the clip, the collision follows the duck: quieter during lines, suddenly present in every gap, still masking on every kick.

    Two generation choices prevent that:

    1. Never write the lead. Prompt or flag the generate as instrumental. No vocal, no vocal to duck around.
    2. Write the lead, then mute it. Split stems. Attach the instrumental (or the pad-without-drums) as the bed. Keep the vocal stem if you need it on a different cut.

    Do this at generate time, not at 11pm in the NLE.

    Order of operations

    1. Decide whether music is the show. If the clip is a chorus-led Short, you want the bounce. Stop. There is nothing to duck because the voice is the track. If the clip is a talking product, a founder line, a tutorial, music is furniture. Furniture does not sing.

    2. Generate the furniture without a lead. In the AI music generator, "instrumental only, no vocals, sparse, leave space in the low-mids for a voiceover" is the brief. "Cinematic, lush, full choir" is how you refill the hole. Length the bed to the picture with extend, do not loop a 12-bar bounce.

    3. If you already have a song you like, split it. Versely's stem split works on in-app Suno generates (separate_music_vocals). ElevenLabs Music v2 has its own 2 / 4 / 6-stem split. Uploaded commercial tracks are a licence problem first and an isolate_audio problem second. Do not stem a record you cannot clear.

    4. Attach the instrumental, not the bounce. Mix mode, not replace. Static music_volume well under the voice. For a continuous VO this is often the whole mix.

    5. Duck only after the lead is gone. Sidechain (or a manual dip) for gaps: cold open, pause, end card. Fast attack, slow release. You are lifting furniture in the holes, not hiding a singer.

    6. Then, and only then, hit a loudness target. Social and streaming have their own LUFS and true-peak ceilings. Those are finishing. If you "master" the stereo bounce first, you have baked the fight into a file that now has no headroom to carve.

    What ducking cannot fix

    Symptom after a duck Actual cause Generation fix
    Hook pops up between sentences Vocal still in the bed Instrumental generate, or mute the vocal stem
    Words feel boxed even when music is quiet Bed occupying 250–500Hz Sparse prompt; high-pass and carve the instrumental, not the VO
    Kick punches holes in the first syllable Drum stem masking speech Four-stem split, pull drums under VO, or generate without a heavy kick
    Music "jumps" mid-clip Loop seam, not a duck Extend the track to length

    If the table row is "vocal still in the bed," no amount of ducking is the next step. Split or regenerate.

    A leaked ghost vocal on an instrumental stem is not a reason to go back to the bounce. It is a reason to listen to the stem solo, once, before you lock. Inferred separators leave tails. If the ghost is loud enough to read as a second speaker, regenerate instrumental. Do not notch a syllable for thirty seconds of ad.

    What not to do

    Do not duck, then stem, then duck again. The timeline now has two music clips and a pair of automation passes that disagree. Stem on the source. One bed clip. One duck.

    Do not treat "music_volume: 20" as stems. That is a static fader on whatever file you attached. If you attached the song, you ducked the song.

    Do not generate a new track to "fix the mix" when the mix is a lead-vs-VO problem. You will get a new lead. Same fight, new key.

    The Versely-shaped version of the good path is one agent request: generate a sparse instrumental, extend it to the video, mix it under the VO. If the generate still sang, split and attach the instrumental only. EQ lives in the NLE. Loudness lives last.

    FAQ

    Can I duck first and stem only if it sounds bad?

    You can. You will do the attach twice. The cheap default is instrumental-or-stems before the first attach, because the first attach is what people "preview" and then forget to replace. Preview a stem. Do not preview a bounce you already know you will throw away.

    Is an instrumental generate always better than a stem split?

    When you know there is VO, yes — nothing to leak. When a vocal version is already the approved cue, split it. Do not regenerate a new instrumental that no longer matches the approved cue's harmony just to avoid a split.

    Does this replace carving 250–500Hz on the bed?

    No. Stems remove the part. Carve removes the band. You still carve an instrumental that was generated as a finished record. You just stop asking a duck to mute a singer. Spectrum work is downstream of this post.

    What if the client sent a finished MP3 as "the track"?

    Ask for an instrumental or stems. If they do not have them, you are mixing a bounce: keep it much quieter than you want, accept the gaps will bloom the hook, and do not promise an M&E. Generating a replacement bed is often cleaner than pretending a stereo master is a stem session.