Speech, voice and audio

    Voiceover

    Also called VO, Narration, Off-camera read.

    A voiceover is a spoken track laid under picture, rather than lipsync that drives a mouth or a caption track that visualises speech.

    The picture already exists, or it will, and the voice is a layer on it. Text-to-speech is how you may create that layer from a script. Voice-cloning is whose voice. Neither of those is the editorial decision to put a read under a clip that does not need a talking face.

    Lipsync is the other door: a mouth has to match. Captions are a third door: the words appear, they are not heard. Dubbing replaces speech that is already in the file. Mixing them is how a product demo gets a generated presenter it did not need, or a VO gets burned in as subtitles with no audio.

    The editing task is add voiceover to video. The generate for the read itself is text-to-speech. Health and finance spots that want a voice without a synthetic presenter use this split on purpose.

    In practice

    • Write the script to time against the picture you already have; do not generate a talking head to carry a read that could sit under B-roll.
    • Generate or clone the voice as its own job, then attach it.
    • If a mouth on screen has to match, that is lipsync, not voiceover.

    The mistake to avoid

    Generating a photoreal presenter because you needed a read. Voiceover does not require a face, and on health or finance a synthetic presenter is the wrong object.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.