Dub with lipsync only when the mouth is the shot
HeyGen-class lipsync is 2.5x credits in the app. Use it when we see the mouth. Voice-only dub when we do not.
Dubbing Studio
One video, three languages
The same clip through Versely's dubbing pipeline: English original, then HeyGen lip-sync into French and Spanish with captions burned in.
Lipsync on a dub is not a quality upgrade you apply by default. It is a second job that rewrites the mouth, and in the app a HeyGen-class lipsync dub is about 2.5× the credits of a voice-only pass. Spend that when the viewer can see the mouth. Skip it when they cannot. A wide of a city, a product turntable, a montage with VO over b-roll — those files need a translated voice, not new teeth.
The chain on this page is the same Dubbing Studio example: English original, then lip-synced French and Spanish with captions. That chain is the expensive path on purpose. It is what you buy when the face is the frame.
Dub a video is the tool. AI dubbing is the product page. This post is the routing rule neither page should have to re-explain: mouth in shot, lipsync. Mouth out, voice-only.
Two engines, two objects
Versely's dub_video job is a localisation pipeline — translate, clone the delivery, write a new soundtrack — not a lipsync model you point at a still. It bills per job, not off the per-second lipsync rate card. Language is not a price dimension. Engine is a capability dimension.
| Engine | What it touches | Length cap | Use when |
|---|---|---|---|
| ElevenLabs (default) | Audio only. Voice cloned into the target language. Picture untouched. | Up to 30 minutes. Trimming supported. | The mouth is not the shot, the file is long, or you need a range inside a longer master. |
| HeyGen | Video. Mouth is re-timed to the new language. | Up to 8 minutes. No in-job trim. | We can see the mouth, and the clip is short enough to fit. |
HeyGen is video-only and supports a smaller language set. ElevenLabs is the default because most files do not need the mouth rewritten. If you are pricing ten markets, ten voice-only jobs is ten times one voice-only job. Ten lipsync jobs is ten times the lipsync job. The 2.5× is per language, not a one-time surcharge for "doing it properly."
A separate tool, generate_lipsync, animates a still into a talking clip. That is not a dub. If you already have footage, do not send it through photo-to-speech to "get lipsync." Send it through dub-video and pick the engine.
When the mouth is the shot
Lipsync earns its credits when a viewer who does not speak the source language would otherwise watch a face that is clearly speaking something else. Talking-head UGC, founder explainers, avatar presenters, interview close-ups, anything framed chest-up with the lips readable at feed size. Those are the files where a voice-only dub looks like a bad foreign mix: the mouth keeps source rhythm, the new language does not.
Even then, lipsync is not magic. It retimes the mouth it is given. It will not hide a profile, a hand over the lips, a bite of food, or a 4K inspect of teeth you never had. If the source performance is a mutter into a collar, HeyGen will mutter in Spanish.
Captions still belong on the dubbed file. Most of the new market will watch muted. Burn them in after the dub, in the target language, not by leaving the English karaoke on the picture.
When the mouth is not the shot
Voice-only is the correct dub for:
- Product and b-roll. Hands, SKUs, rooms, motion graphics. There is no mouth to save.
- Films that cut away. A film someone actually pays for is usually a mix of talking coverage and coverage that is not talking. Lipsync the close-ups if you must. Do not lipsync the reel.
- VO-led explainers where the speaker is never on camera, or is a silhouette, or is out of focus on purpose.
- Masters longer than eight minutes. HeyGen will not take them. Cut the talking section out if it needs lipsync; leave the rest on ElevenLabs.
The expensive mistake is sending the whole timeline through HeyGen because "localisation should look native." Native, for a pack shot, is a clean translated VO that does not fight the picture. Native is not a mouth that was never there.
Approve the master, then fan out
Every resubmission is another job at the same price. Get the English (or source) cut signed before you multiply it. Trim dead air on the default engine inside the job; HeyGen will not trim for you, so cut the file first.
Do not mix engines on one market "to see." Pick by shot type, then run the batch. If a campaign is half talking-head and half product, split the timeline: lipsync the A-roll, voice-only the rest, recut. That is cheaper than lipsyncing a three-minute film because thirty seconds of it has a face.
Source has to already live on Versely — a generation, an upload, an output. A YouTube URL is not a dub source.
A useful internal rule: if you would be embarrassed to ship the file with the sound off because the mouth is the only performance, pay for lipsync. If the file still makes sense as a silent product demo, do not. Credits follow the face, not the ambition of the localisation plan.
FAQ
Is lipsync always 2.5×?
In the app, a HeyGen-class lipsync dub lands around two and a half times a voice-only dub of the same file. Confirm with the in-app estimate before a ten-language batch. Do not price it off the per-second lipsync catalogue; a dub is a different job type.
Can I lipsync only the talking section of a longer video?
Yes, by cutting that section first. HeyGen will not take a start/end range and caps at eight minutes. ElevenLabs will take a range on files up to thirty minutes, without rewriting the mouth.
Does a dubbed talking head still need captions?
Yes. Lipsync is for people who can see the mouth and hear the file. The muted feed still needs type, in the language you dubbed into. Caption after the dub, not before — otherwise you burn source-language words onto the new voice.
Should I use generate_lipsync instead of dub_video?
Only when the input is a still you want to become a talking clip. Existing footage that needs a new language is dub_video. Photo-to-speech will not preserve the original performance; it will replace it.