Dubbing a video clones a voice into a language; captions do not
dub_video is a cloned-voice language job on Versely-hosted media. Subtitles are a different job, and using the wrong one wastes credits.
One video. A new language, in a cloned voice. That is the job. Dub a video into another language is one named agent capability. dub_video is full AI dubbing — voice cloning, optional lip-sync — not a subtitle burn and not a TTS read of a translation you typed. Using it as captions wastes a dubbing job on a file that only needed typeset speech. Using captions as a dub wastes a shipping week on an audience that cannot hear the language.
AI dubbing is the spoken replacement. Captions are the written overlay. They are not a fallback for each other.
What dub_video will actually run
Give the agent a Versely-hosted source_url (upload or prior generation — not a YouTube link) and a target_lang. The call submits an async job and returns a project id immediately. Poll get_dub_status or ask the agent to check list_dubs. Do not sit on the turn waiting for minutes of audio to finish.
Two engines, two contracts:
- ElevenLabs (default): audio-cloning dub, works on audio or video, optional trim via
start_time/end_time, up to 30 minutes. - HeyGen: video-only, lip-synced translation, up to 8 minutes, no trimming, and a smaller language set.
Pick the engine for the constraint you actually have — on-camera mouth, runtime, language — not for the brand name you remember. Credit cost is shown before the job starts. The AI dubbing tool and dub video are the same work with a form. This page is the agent verb.
What this job is not
Subtitles. Transcribe and caption my video burns word-timed captions from real speech. The viewer still hears the original language. If that is enough, do not dub. Dubbing a clip so you can "also have English on screen" is a cloned performance you will then caption anyway.
A text overlay you wrote. Add a text overlay to my video stamps one authored line. It does not listen. It does not translate.
A new talking-head generate in the target language. Make a talking avatar video animates a face to audio you already have. Dubbing starts from a finished video or audio file and replaces the language. Generating a second presenter "in Spanish" is a new identity, not a dub of the first.
An external URL. The tool's own limit: source must already be hosted on Versely. Pasting a public YouTube link into this job does not fetch it. Upload or generate first, then dub.
Dub the keeper. Caption the dub if you still need type.
The waste is dubbing a cut you are still rewriting. ElevenLabs will clone whatever performance is on the file, including the ums you were about to trim. HeyGen will lip-sync that same unfinished take, up to eight minutes, and you will pay to do it again after the trim.
Lock picture and original audio. Confirm the language and the engine. Submit. While it runs, do not also generate a "Spanish version" on a talking-head model "in case the dub is slow." That is a second film.
If the audience needs both a new language and on-screen words, dub first, then caption the dubbed file. Captioning the original and hoping the type is the localization is how you ship English burnt over a Spanish track, or the reverse.
FAQ
Can I dub a YouTube link directly?
No. source_url must already be Versely-hosted. Upload the file, or dub a generation you made here. An external link is not a source this tool will fetch.
Is HeyGen always the right engine because of lipsync?
Only when the speaker is on camera in close-up, the runtime is under 8 minutes, and the language is in HeyGen's set. Audio-only, long-form, and trim-a-section jobs belong on ElevenLabs. Forcing lipsync on a podcast wastes a video engine on a file with no mouth.
Do I lose the original if I dub?
No. Dubbing returns a new file (or audio). The source stays. That is why you dub a keeper instead of overwriting it in place with a generate.
Why not just burn translated captions and skip the dub?
If the audience reads and will tolerate the original audio, captions are the cheaper named job. If they need to hear the language, captions are not a dub. Run dub_video. Do not ask a caption tool to clone a voice; it cannot.