It is a chain, not a single operation: transcribe the original, translate it, speak the translation in a voice matched to the original speaker, then fit the result back to the picture. Each stage can fail independently, and the most common complaint about a dub traces back to the translation rather than to anything about the audio.
Timing is the hard constraint nobody anticipates. Languages take different amounts of time to say the same thing, so a faithful translation frequently does not fit the gap where the original line sat. Something has to give — the translation gets tightened, the delivery gets compressed, or the fit gets loose — and which one a tool picks is most of why some dubs feel natural and others do not.
Lipsync is the optional last stage. Without it you have a well-fitted voiceover; with it, the mouth agrees with the new language too.
In practice
- Review the translation before the audio — that is where most quality is won or lost.
- Idioms, brand names and numbers are the usual errors; check them explicitly.
- Keeping the original speaker's voice is what makes a dub feel like the same person, not a stand-in.
The mistake to avoid
Dubbing a video whose on-screen text is still in the original language. The audio switches, the graphics do not, and the mismatch is immediately obvious.
Go deeper
Inworld TTS-2 for dubs, then a lipsync model
Inworld TTS-2 is the high-volume dub voice, priced under ElevenLabs v3. ElevenLabs v3 is the premium read. Sync.so or Hedra is the lipsync model after that.
Where you will run into it
- Dub a Video Into Another Language — Same voice, same face, new language.
- Translate a Video Into Another Language — Just the words, translated — lighter than a full dub.
- AI Dubbing Tool — One approved video. Several markets.
- AI Lipsync Generator — Text, audio or video in. Talking head out.
Related terms
Lipsync
Lipsync meaning: driving a face mouth from an audio track so speech reads as spoken, not dubbed over the top.
Voice cloning
Voice cloning meaning: building a reusable synthetic voice from a real sample so new scripts can be spoken in that voice later.
Speech-to-text
Speech-to-text meaning: converting spoken audio into written text, the transcript captions, translation, and search all depend on.
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
Forced alignment
Forced alignment meaning: matching a known transcript to its audio to find exactly when each word was spoken.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free account. Works in your browser - no install needed. The same account signs in on your phone.