Comparisons

    Two Dubbing Engines and Their Hard Limits

    Choosing a dubbing engine looks like a quality preference. It is really a constraint problem: runtime, lip movement, trimming, and language coverage.

    Versely Team7 min read

    Dubbing engine choice usually gets talked about like a taste question — which one sounds more natural, which lip-sync looks less uncanny. That conversation is premature more often than not. Behind Versely's dubbing tool sit two genuinely different engines, and in most real jobs the source file itself eliminates one of them before quality ever gets a vote.

    The two engines are different transformations, not different tiers

    ElevenLabs, the default, is an audio-cloning dub: it clones the source voice and generates translated audio, and it works on either an audio file or a video file. It doesn't attempt to re-sync lip movement — it swaps the audio track and leaves the video as-is.

    HeyGen is video-only, lip-synced translation: it doesn't just translate the audio, it re-times the mouth movement in the video itself to match the new language.

    That's not a quality gap between a "better" and "worse" option. It's two different jobs. One changes what a viewer hears. The other changes what a viewer sees and hears.

    Four constraints that decide it before quality does

    • Source type. An audio-only source — a podcast, a voiceover track, a music bed with vocals — is an ElevenLabs job by necessity. HeyGen is video-only; there's no footage for it to lip-sync.
    • Runtime. ElevenLabs runs up to 30 minutes; HeyGen caps at 8. A 12-minute explainer is an ElevenLabs job regardless of any lip-sync preference, because HeyGen simply won't accept it.
    • Trimming. ElevenLabs supports dubbing only a portion of a longer source via start and end times. HeyGen doesn't support trimming at all — a HeyGen job dubs the entire source, full stop. If the goal is localizing a 90-second segment out of a 6-minute video, that's either a pre-trimmed source or an ElevenLabs job.
    • Language coverage. ElevenLabs supports the broader language set; HeyGen supports only a subset. A target language can rule out HeyGen entirely before lip-sync quality is even relevant — worth confirming before planning a lip-synced campaign across a long language list, not after.

    Multi-speaker sources get a shared setting too — a speaker-count hint applies to the dub job regardless of which engine ends up running it, which matters for panel or interview-style footage with more than one voice to keep distinct.

    The decision in practice

    Run it as a sequence, not a preference:

    1. Is the source audio-only? → ElevenLabs.
    2. Is the source over 8 minutes? → ElevenLabs.
    3. Does the job need only part of a longer file dubbed? → ElevenLabs, using the trim window.
    4. Is the target language in HeyGen's supported set, and does the piece live or die on a face close to camera where mismatched lip movement would be the most visible flaw? → HeyGen, once it's cleared the first three gates.

    In practice, most long-form and most audio-first content is an ElevenLabs job by elimination, before anyone's compared how either engine actually sounds. HeyGen earns its slot specifically for shorter, face-forward video where an obviously-dubbed mouth is the flaw a viewer would notice first.

    What "quality" means once you're through the gates

    Only after the constraints have already picked a lane does quality become the differentiator. ElevenLabs' strength is natural cloned-voice audio plus flexibility — any source type, longer runtime, partial dubs on demand. HeyGen's strength is specific to the moment a viewer is staring at a mouth and the improved lip-sync is the entire point of the localized cut. Neither engine is "better" in the abstract. They're solving different visible problems, and the job in front of you usually only has one of those problems.

    The knobs that matter after you've picked an engine

    A handful of settings apply once the engine is decided, and they're worth knowing before a job runs rather than after the output looks wrong:

    • translate_audio_only skips the video track's other processing and outputs just the translated audio — useful when the plan is to re-cut or re-time the dub manually rather than take the tool's assembled output as final.
    • drop_background_audio controls whether music and ambience under the original speech survive the dub or get stripped along with it. Leaving background audio in matters most for content scored with music the dub shouldn't silence; dropping it matters when the source has noise or room tone that would sound wrong carried into a cleanly-cloned voice track.
    • highest_resolution and watermark are output-quality and branding toggles worth checking against whatever the destination platform actually needs, rather than leaving on default and finding out after export.
    • enable_dynamic_duration lets the output run longer or shorter than the source to accommodate a translation that doesn't fit the original timing — German famously runs longer than English for the same meaning, and a dub locked to the exact original runtime either speeds up unnaturally or clips the end of a sentence. Turning this on trades a fixed runtime for a natural pace, which is usually the right trade for anything that isn't cut to an exact ad-length slot.

    None of these decide which engine to use — the four constraints above do that first. They decide how the chosen engine's output actually sounds and fits once the harder decision is already made.

    A Versely walkthrough

    Two real jobs, run through the sequence above:

    A 12-minute podcast episode, audio only, needs Spanish, French, and German — but only from the intro through the 8-minute mark, skipping a sponsor read at the end.

    Prompt: "Dub this episode into Spanish, French, and German, from the start through 8:00, and keep the background music."

    This is ElevenLabs by three separate constraints at once: it's audio, it's over 8 minutes, and it needs a trim window — HeyGen fails all three gates before language or voice quality is even a conversation.

    A 4-minute founder talking-head video for a landing page needs a Japanese version with real lip movement, not a voice-over mismatch.

    Prompt: "Dub this founder video into Japanese with lip-sync."

    Under 8 minutes, video source, and lip-sync is the explicit ask — a HeyGen job, provided Japanese is inside its supported language set, which is worth checking against Versely's dub-a-video capability page before committing to it across a longer language list.

    Both jobs run asynchronously — the agent submits the dub and hands back a project ID to poll rather than a same-turn result, so "is my dub done yet" is a follow-up question, not something to wait on mid-conversation. And whatever gets dubbed has to already be hosted on Versely — a prior generation, upload, or output — since an external link isn't a valid source yet; downloading and re-uploading outside footage is step zero, not an afterthought.

    One structural note worth knowing before dispatching a multi-language batch: dubbing bills as one flat job per submitted language, not by the minute and not by which language you picked — the tenth language costs the same as the first, and the only thing that multiplies the total is how many target languages you're running. The dubbing-into-10-languages cost breakdown walks through that shape in full, and it's a large part of why engine choice deserves more thought than language choice does: the engine changes what's technically possible, where the language list mostly just changes the job count.

    For everything else that sits around a dub — script prep, voice selection, and the wider localization surface — Versely's dub-video task and the voice-over hub are the two pages worth bookmarking alongside this one.