Guides

    Write the words: generate_speech speaks them, it does not finish the video

    Write and generate a voiceover is one agent job. generate_speech returns a spoken file. Laying it on picture, or casting two speakers, is a different job.

    Versely Team3 min read

    Write the words. generate_speech speaks them. Write and generate a voiceover is that job and only that job. It is not attach-to-picture, not a two-host podcast, not a dub. Spending a speech generation on a script you will throw away when the cut changes, or chaining several solo reads because you needed a conversation, is how TTS credits disappear without a usable soundtrack.

    The tool is generate_speech. Required: the script text and a voice model. Optional: voice_id, emotion, style_instructions, language. You get an audio clip of the script spoken aloud. The text-to-speech tool is the same read with a different door.

    A read is not a lay

    The output is audio. Picture does not move. Mouths do not lock. If you already have a video URL and a track you like, add music or a voiceover to a video is the lossless attach (replace or mix). Generate first only when you still need the take.

    A warm, confident narrator on a 12-second product line is this page. "Read this in Spanish, calm" is this page. "Put this VO on the clip and duck the room" is attach. "Alex and Jamie riff for 40 seconds" is create a multi-voice dialogue or podcast, which is generate_multi_speaker_speech in one call. Hand-stitching three generate_speech files into a fake conversation is a timing puzzle you will pay to rebuild.

    Picture lock before a hero read. A new take has a new duration. The speech file will overhang, die early, or describe a shot you cut. Scratch VO on a maybe-clip is fine if you treat it as scratch and do not pay for a hero model on a file you already distrust.

    Steer the delivery, not a second job

    emotion and style_instructions change the performance: warm and confident, fast and energetic, a flat read if you leave them blank. That is still one speaker saying your text. It does not write the script for a series, clone a voice (that is a separate sample-in job), or translate a finished video. Dubbing a talking clip is not TTS.

    Cost is per voice model, shown before you confirm. Confirm the words and the voice, not the whole video.

    FAQ

    Does generate_speech attach the voiceover to my video?

    No. It returns audio. Attach audio is the lay, with replace or mix. Running speech again hoping the pictures change wastes a second read.

    When do I use multi-speaker instead of several generate_speech calls?

    Whenever there are two named voices in one conversation. One tool, one file, lines in order. Chaining solo reads is how a draft interview becomes a stitch you redo on every recut.

    Can I generate the VO before the picture exists?

    You can. Treat it as scratch unless duration is already locked. A keeper that is eight seconds will not carry an eighteen-second script. Shorten the copy on the keeper, then generate the real take.

    Is a cloned voice this job?

    No. Clone from a sample you own, then pass that voice_id into generate_speech. This page speaks a script. It does not mint the voice.