It is a chain, not a single operation: transcribe the original, translate it, speak the translation in a voice matched to the original speaker, then fit the result back to the picture. Each stage can fail independently, and the most common complaint about a dub traces back to the translation rather than to anything about the audio.
Timing is the hard constraint nobody anticipates. Languages take different amounts of time to say the same thing, so a faithful translation frequently does not fit the gap where the original line sat. Something has to give — the translation gets tightened, the delivery gets compressed, or the fit gets loose — and which one a tool picks is most of why some dubs feel natural and others do not.
Lipsync is the optional last stage. Without it you have a well-fitted voiceover; with it, the mouth agrees with the new language too.
In practice
- Review the translation before the audio — that is where most quality is won or lost.
- Idioms, brand names and numbers are the usual errors; check them explicitly.
- Keeping the original speaker's voice is what makes a dub feel like the same person, not a stand-in.
The mistake to avoid
Dubbing a video whose on-screen text is still in the original language. The audio switches, the graphics do not, and the mismatch is immediately obvious.
Where you will run into it
- Dub a Video Into Another Language — Same voice, same face, new language.
- Translate a Video Into Another Language — Just the words, translated — lighter than a full dub.
Related terms
Lipsync
Lipsync generation drives a face's mouth from an audio track, so the speech reads as spoken rather than dubbed over the top.
Voice cloning
Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.
Speech-to-text
Speech-to-text converts spoken audio into written text, producing the transcript that captions, translation and search all depend on.
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
Forced alignment
Forced alignment matches a known transcript to the audio it came from, working out exactly when each word was spoken.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.