Workflows

    Cut M&E Stems Before You Need a Dub

    Keep DX, MX, FX and optional ambience as separate files, with a sync pop, so a later-language dub is a voice swap, not a remix.

    Versely Team9 min read

    A baked stereo mix is a finished record, not a localization master. The moment music and effects share a file with the voice, the next market is a remix: someone has to pull the English (or whatever you shot) out of a bed that was never meant to let go of it. Cut the stems before anyone asks for a dub, and the later-language job is a voice swap against files you already have.

    This is a delivery spec: dialogue (DX), music (MX), effects (FX), optional ambience, and a sync pop. If any of those live only inside a combined print, you do not have an M&E. You have a hope that isolation will invent one.

    Why a baked mix kills the next market

    AI dubbing replaces speech. It does not reconstruct a music-and-effects track that was never saved. Versely's dub tool clones source speech into a target language (and, on one engine, retimes mouths). The two engines decide runtime, trimming, and whether lip movement is even available. Neither engine can un-bake a pad that was generated into the same waveform as the presenter.

    drop_background_audio is the salvage knob, not the stem session. Leave background in and the new voice sits on the original bed, which is fine only if that bed contains no original-language bleed. Drop background and you throw away the score and the spot effects with the old voice. Both outcomes are what you get when DX, MX, and FX were never separate.

    A native-audio generate makes this worse by construction. Native audio writes picture and sound in one pass. Room, score, whoosh, and line arrive as a single stereo file. That is the right product for a one-language social cut. It is the wrong printmaster for a campaign that will need Spanish in six weeks.

    Isolation is not an M&E either. Splitting a generated Suno track returns vocal and instrumental halves of a song you made in-app. isolate_audio pulls speech out of a mixed clip you point at with a URL. Both are two-way splits. A localization stem session is intentional tracks, generated as themselves, not inferred from a bounce.

    The delivery spec: DX, MX, FX, ambience, sync pop

    Name the files the way a mixer already does. Do not invent a house code that only your chat log understands.

    Stem What lives on it What must not live on it
    DX Dialogue, VO, production speech, breaths that belong to the line Score, spot FX, room tone you intend to reuse under a dub
    MX Score and source music Sung hooks fighting the VO, unless that hook is the product
    FX Sync effects: hits, whooshes, UI, footsteps you placed Continuous room, music, speech
    AMB (optional) Presence, fill, air, crowd bed Discrete events that have to hit a picture cut
    Sync pop One-frame 1 kHz tone, two seconds before first frame of picture Programme audio

    Three stems are required: DX, MX, and FX. Ambience is a fourth when the room has to survive a language change. If you skip AMB, say so in the delivery note. The sync pop is a leader on every file, not a mix stem of its own.

    The sync pop is editorial glue. A one-frame 1 kHz tone two seconds before action (the conventional 2-pop, with a matching tail pop after last frame if the job will be conformed) lets a later session line up picture and stems. Versely's default video frame rate is 25 fps. Write the pop against that rate unless the spec names another, and do not mix 24 fps pops into a 25 fps sequence.

    Bounce each stem as a full-length file, even the quiet ones. A 12-second FX file against a 40-second DX file is how sync gets lost in a zip. Same start, same duration, same sample rate, same pop.

    A printmaster (the full mix) can sit next to the stems. It is a reference, not a substitute. Anyone who only keeps the printmaster has already decided the next market will pay for a remix.

    How generated video bakes the problem in

    Native-audio video. The model scores the room, the line, and often a hint of music because "cinematic" was in the prompt. There is no DX file. If the shot is the one you will ship, generate the line again as speech, generate the bed and the FX as themselves, and rebuild.

    Added audio, mixed too early. You generated VO, music, and SFX as separate jobs, mixed them, and exported once. The MP4 is the printmaster. If the source WAVs were not kept, you are back to isolation. attach_audio_to_video is a bounce, not an archive.

    A dub of a baked master. If you already chose dub, the input should be DX-replaceable. Feeding the tool a social export with a loud bed forces drop_background_audio into a lose-lose: keep the bed and risk original-language ghosts, or drop it and relight the whole soundtrack.

    The failure is the moment someone exported a mix and deleted the inputs.

    Build the stems as you generate

    Keep the layers apart from the first generate, not as a cleanup pass.

    1. Picture silent, or picture with native audio you will throw away. If the clip must localize, do not spend prompt effort on a score inside the video model. Lock picture. Replace audio later.
    2. DX from speech, not from the clip. Write the script and generate it with text-to-speech (or a clone, if that is the voice you actually have rights to). That file is the DX stem. It can be re-generated in another language without touching MX or FX.
    3. MX as an instrumental bed. Prompt the music generator for instrumental, sparse enough to sit under speech. If a vocal version exists, split it and archive the instrumental as MX. Do not use the full mix as the bed.
    4. FX as their own generates. Generate the sound effect for the hit, the whoosh, the UI chirp. Place them on a dedicated track in the NLE. Do not ask the music model to "add cinematic hits."
    5. AMB only if the room is a character. A cafe bed that has to continue under a dubbed line belongs on AMB, not under DX. If the VO is already dry, you can add presence once and reuse it in every language.
    6. Pop, then bounce. Same duration on DX, MX, FX, optional AMB, and the printmaster, with the same 2-pop on every file: job_dx.wav, job_mx.wav, job_fx.wav, job_amb.wav if used, job_printmaster.wav.

    Keep that folder, a language list, and a note that DX is the stem a dub should replace. MX and FX do not get re-generated per market unless the brief changes the music.

    If you assemble in Versely's editor, treat the export as the printmaster and keep the inputs. A 480p preview pass (preview: true) is free, with a short per-user cooldown between passes, and it is the right place to judge whether the voice sits. It is not an archive. The WAVs you generated are the archive.

    Salvage paths, and what they cannot restore

    Sometimes the baked file is all you have. Be honest about what each tool returns.

    Starting point Tool What you get What you still do not have
    In-app Suno track separate_music_vocals Vocal + instrumental DX/FX/AMB, a sync pop, a picture-locked M&E
    Any mixed clip isolate_audio A voice-ish file with bleed A clean MX, placed FX, matching duration stems
    Baked video, new language dub_video with drop_background_audio on or off New speech, with or without the old bed Discrete MX/FX you can remix
    Baked video, you only need new VO Replace the audio with a new TTS line A new printmaster Any later market without repeating this

    Stem separation on a song is a two-way split with leakage. Use it to mute a singer under a VO, not to manufacture a film M&E. Isolation removes what is around a voice. It does not give you a whoosh on frame 17.

    If the job is one language and will stay that way, a printmaster is enough. Cut stems on day one, while the generates still exist as separate URLs.

    FAQ

    Do I need an M&E if I am only publishing in one language?

    No. A printmaster is enough for a single-market cut. Cut stems when a second language is plausible, a client asked for localizable deliverables, or the same picture will be recut with a different VO. Extra bounces now beat a remix later.

    Can I rebuild stems from a native-audio clip?

    Not as stems. You can replace the whole soundtrack: new TTS for DX, new instrumental for MX, new SFX for FX, then mix. The original native-audio file is a reference for timing, not a source you split. If the picture depends on in-pass mouth movement, you are choosing between a dubbed mouth and a regenerate in-language.

    Where does the sync pop go if the video already starts at timecode zero?

    Add two seconds of pre-roll. Pop at the start of that pre-roll (one frame of 1 kHz), first frame of picture two seconds later, optional tail pop two seconds after last frame. Bounce every stem with that same pre-roll. A pop on top of programme audio is how the first word gets clipped in the next session.

    Is drop_background_audio an M&E button?

    No. It is a binary on the dub job: keep the source bed or strip it. An M&E is a mastered music-and-effects track built without dialogue, with FX and music in their own files so a new DX can sit down without a fight. If you need that, you cut it before the dub, not inside it.