Guides

    Narration is a finishing track, not a way to hide a miss

    generate_speech plus attach_audio_to_video lays a script on a clip. Narrating a take you will regen, or narrating before picture lock, pays for a read you will orphan.

    Versely Team3 min read

    A voiceover is a script that became audio, then a clip that received it. generate_speech is the read. attach_audio_to_video is the lay. Decide replace or mix up front: replace if the original sound is trash, mix if you are narrating over B-roll room. Either way, the picture has to be the picture. A new take has a new duration. The read will overhang, or die early, or talk over a shot that no longer exists.

    Add a voiceover to a video is the job. Billing is per speech generation (model and script length) plus a flat attach fee. estimate_cost before a long script. The TTS tool is the same read in a different door.

    The read does not follow the picture

    generate_speech starts at the beginning. It is not auto-spotted to on-screen actions. If you needed a line at 0:04 because the product enters at 0:04, you still have to build a picture whose 0:04 is stable. Recutting after the read is how you get a narrator describing a shot you cut.

    People record VO first because the script feels like the brief. In a generated pipeline the brief is not the file. The file is the file. Write the script. Sample picture until the sequence can carry it. Then read. If the keeper is 8 seconds and the script is 18, you needed a shorter script or a longer picture — decided on the keeper, not in the abstract.

    Emotion and style_instructions vary by TTS model. Do not spend a "warm, conversational" pass on a clip you already know is the wrong room. The performance is not transferable. The words might be. Paste them again onto the next file when that file is real.

    One voice versus a cast

    A single narrator is this task. Two named speakers is generate_multi_speaker_speech. Chaining several generate_speech calls and hand-stitching them is how a draft podcast becomes a timing puzzle you will redo when the picture changes. If you even might recut, do not build the stitch.

    Clone the house voice once, if you need identity, from a clean sample you own. Then use that voice_id on keepers. Cloning is not a reason to narrate the whole bin.

    Order

    1. Lock picture and duration, or accept that the VO is scratch and will be thrown out (do not pay hero-model rates for scratch).
    2. Generate the read.
    3. Attach with the mode you meant. Replace is mute-the-original. Mix keeps room.

    If you are still choosing between two product orbits, neither of them gets the hero VO.

    FAQ

    Can I reuse one speech file across several takes?

    Only if duration and in-point match. A 9-second read on an 8-second regen will clip or leave a hole. Re-generate or trim the audio against the keeper. Do not pretend attach is elastic.

    Mix or replace for B-roll with faint room?

    Mix, with VO louder than the room. Replace if the room is noise. Decide on the keeper; a miss with ugly room is not a mix problem, it is a take problem.

    Should I caption before or after the voiceover?

    After. Captions transcribe this soundtrack. VO first, on the locked picture, then burn captions. Captioning the silent generate is a blank track you will pay to do again.

    Is a long script cheaper to test on a still?

    Yes. Read it out loud yourself. Pay generate_speech when the sequence can hold the sentences.