The picture already exists, or it will, and the voice is a layer on it. Text-to-speech is how you may create that layer from a script. Voice-cloning is whose voice. Neither of those is the editorial decision to put a read under a clip that does not need a talking face.
Lipsync is the other door: a mouth has to match. Captions are a third door: the words appear, they are not heard. Dubbing replaces speech that is already in the file. Mixing them is how a product demo gets a generated presenter it did not need, or a VO gets burned in as subtitles with no audio.
The editing task is add voiceover to video. The generate for the read itself is text-to-speech. Health and finance spots that want a voice without a synthetic presenter use this split on purpose.
In practice
- Write the script to time against the picture you already have; do not generate a talking head to carry a read that could sit under B-roll.
- Generate or clone the voice as its own job, then attach it.
- If a mouth on screen has to match, that is lipsync, not voiceover.
The mistake to avoid
Generating a photoreal presenter because you needed a read. Voiceover does not require a face, and on health or finance a synthetic presenter is the wrong object.
Where you will run into it
- Add a Voiceover to a Video — Type the script. Get a narrated video back.
- AI Text to Speech — One script, several engines, one bill in credits.
Related terms
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
Lipsync
Lipsync generation drives a face's mouth from an audio track, so the speech reads as spoken rather than dubbed over the top.
AI dubbing
AI dubbing replaces a video's spoken audio with another language, usually keeping the original speaker's voice and optionally re-syncing their mouth.
Burned-in captions
Burned-in captions are subtitles rendered into the video's pixels, so they cannot be switched off, restyled by the player, or lost when the file is re-uploaded somewhere else.
Voice cloning
Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.