Workflows

    Multilingual Lip Sync Without a Separate Dubbing Pass

    Generating dialogue natively in the target language can replace the dub-and-lipsync pipeline for net-new creative — but not always. Here's the split.

    Versely Team7 min read

    The default assumption for years has been fixed: shoot or generate the video once, then dub it into every other language afterward. That order made sense when dubbing was the only tool that existed. It's no longer the only order available. A video model that generates dialogue natively can write and speak the line in Spanish, German, or Japanese at render time — no English master, nothing to translate away from, no mouth movements shaped for one language and reworked for another. For content that doesn't exist yet, that changes which pipeline is actually the right default.

    It doesn't replace dubbing outright, though, and knowing which of the two to reach for is the actual skill here.

    World map with connection points across continents

    The two pipelines, plainly

    Dub-first is the established path: produce one master video, then translate and re-voice it per market, re-syncing the mouth to match. It's what Versely's dub video tool does, and the full workflow for taking one master asset across markets is covered in depth elsewhere on this blog — voice cloning, engine choice, and the localization QA that makes it hold up.

    Generate-native is the newer option: skip the master entirely, and produce an independent clip per target language, with the dialogue written and spoken in that language from the first generation. There's no original take being translated, because there was never a take in a different language to begin with — each version was native from the start.

    Why generate-native can beat dubbing for new creative

    The quality argument comes down to what the mouth movement is actually derived from. In a dubbed clip, the lip-sync pass has to reshape mouth movement that was originally performed or generated for a different language's rhythm and syllable count, retrofitting it to match translated audio after the fact — a harder problem than it sounds, since languages don't share timing. A native generation doesn't have that retrofit step: the model is producing the mouth shapes and the dialogue audio for that language in the same pass, so the two agree by construction rather than by correction.

    Flux 3 Video is built specifically around this. Its own documentation lists support for English dialects, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi and more, with precise lip-syncing across all of them — meaning the sync is a property of how each language version was generated, not something bolted on to a video that was never speaking that language in the first place.

    Generate-native also solves a pacing problem dubbing can't: languages don't run the same length for the same idea. A German sentence is routinely longer than its English equivalent; a Japanese line may compress differently again. A dubbed clip is stuck with the original footage's timing regardless of what the translation actually needs, which is why dub tracks so often sound rushed or padded. A native generation paces the whole shot — cuts, on-screen text, dialogue — to what that language's version actually requires, because nothing about the shot was locked before the language was decided.

    Where dubbing still wins

    None of that makes generate-native the right call universally, and the honest cases where dub-first still wins are worth being specific about:

    • The subject is a real person. A founder's testimonial, a real customer, a signed talent's on-camera read — you cannot regenerate a real human speaking a new language. You can only take their real footage and lip-sync the mouth to match a new audio track, which is exactly what a dedicated dub pipeline is for.
    • An expensive, approved edit already exists. If a hero asset went through a full production and approval cycle — specific pacing, specific B-roll, a locked cut — dubbing preserves that one approved edit across languages instead of asking a client to re-approve N independently generated versions.
    • The copy is regulated or legally exact. When a translation has to be verified word-for-word by a professional translator, generating a script and dubbing it in gives you direct control over the exact translated text. Trusting a generation model's own multilingual output to match an approved legal translation exactly is a harder guarantee to make.

    The rule that falls out of this is simple: if the subject is real, dub it. If the subject is synthetic and the goal is several independently native-feeling versions, generate each one directly in its language instead.

    Running both paths in Versely

    For net-new creative with no existing footage to protect, native generation is one request per language, run in parallel rather than in sequence:

    "Generate this 15-second product ad on Flux 3 Video in Spanish, German, and Japanese — same visual brief and same product shots, but write and deliver the dialogue natively in each language rather than translating a single script."

    Each version comes back independently paced and independently lip-synced, with nothing to reconcile against an English original because there wasn't one.

    For the founder-testimonial case, where the footage is real and has to stay real, the dub path runs through the source clip instead:

    "Dub our founder's testimonial into Spanish and Japanese using the lip-synced video engine, so the mouth movement matches the translated audio on the real footage."

    That dispatches against the actual video, translating and re-syncing rather than regenerating the speaker from scratch — the right call whenever the person on screen has to remain recognizably the same person. Versely's catalog carries the narrower version of this too: if a clip already has translated audio ready and just needs its mouth resynced to match, lipsync video handles that resync alone without the translation and voice-cloning steps a full dub pass includes — useful when the audio side of the job is already solved and only the mouth needs to catch up. Underneath both paths sit dedicated models built for exactly this kind of resync work, including Kling Lipsync for matching an existing clip's mouth to new audio and Wan 2.2 Speech to Video and LTX 2.3 Audio to Video for audio-driven lip animation — the machinery a dub pipeline runs on, as distinct from a model generating a fresh performance from scratch.

    Whichever path fits the brief, the best lipsync model ranking is worth checking before committing to an engine — quality varies meaningfully by language pair, and the model that handles a Spanish resync best isn't guaranteed to be the same one that handles Japanese best. For the dub path specifically, casting or cloning the voice that carries the translation happens in the voice-over toolset before the dub step ever runs.

    FAQ

    Do I need a different script for native generation than for dubbing?

    Yes, in spirit if not in tooling. A dub script is a direct translation constrained to match the original footage's timing. A native-generation script can be written to fit the target language's natural pacing from the start, since there's no existing footage timeline it has to squeeze into.

    Can I use both approaches in the same campaign?

    Regularly, and it's often the right call — a real founder or testimonial segment gets dubbed to preserve the actual person, while fully synthetic product or B-roll segments in the same campaign get generated natively per language for better pacing and sync.

    Which pipeline is faster for a brand-new multilingual campaign?

    Generate-native skips the translate-and-resync step entirely for content that doesn't exist yet, which removes a stage from the pipeline. Dubbing is faster when a master asset already exists and re-approving a full new edit per language isn't necessary or wanted.

    What if the target language isn't well supported by a native-generation model?

    Fall back to the dub pipeline for that language specifically — dubbing engines generally cover a broader language list than any single native-generation model does, so mixing approaches per language within one campaign is normal rather than a sign something went wrong.