Guides

    Cast a multi-voice track against a locked cut, not a draft

    generate_multi_speaker_speech builds a conversation in one call. Attaching it before picture lock orphans a timed dialogue on a recut.

    Versely Team3 min read

    A two-hander is one audio file, not two glued VO jobs. generate_multi_speaker_speech assigns a distinct voice to each speaker alias in a single call. attach_audio_to_video lays the result on a clip. That conversation has internal timing. Recut the picture and the Guest answers a question the Host no longer asks on screen. You will regenerate the whole cast.

    Add multi-speaker dialogue to a video is the job. Billing follows script length and speaker count. estimate_cost before a long conversation. A single narrator is still generate_speech. Reach for the cast only when you actually need two or more voices.

    The timing is the asset

    Each alias in the script needs a matching entry in speakers. Miss that and the assignment is ambiguous. Get it right and you still have a track whose pauses assume a picture. Podcast-style clips, interview ads, scripted banter — all of them are spotted against this cut. A draft sequence with placeholder B-roll is not a cut.

    People generate the dialogue first because the script is fun. Then they generate picture to "cover" it. Then the picture is 20% shorter. Then they regenerate the conversation. The first generate was a table read you paid model rates for. Table-read in a doc. Pay the tool when the cut can hold the lines.

    Attaching a multi-speaker track to a miss also burns the attach fee. The speech file might be reusable if duration matches a later keeper. Duration rarely matches. Plan to regenerate the conversation against the locked file, which means do not generate it yet.

    Visuals are a separate later step

    This tool returns audio. On-screen speakers, split-screen, lipsync per person — none of that is included. Driving two portraits with generate_lipsync against the same conversation is two more models, each of which needs locked stills and this locked track. Stacking those on a draft is how a "quick podcast ad" becomes four finishing passes on files you will replace.

    The AI video generator still has to produce plates worth covering. Dialogue is not coverage.

    Order

    1. Lock picture duration, or accept MOS plates with a duration you will not change.
    2. Generate the multi-speaker track with aliases that match speakers.
    3. Attach with replace or mix, on purpose.
    4. Caption the conversation, not the silent picture you had yesterday.

    FAQ

    Can I chain two generate_speech calls instead?

    You can, and you will hand-stitch timing yourself. When the picture changes you restitch. One multi-speaker call is the right tool for 2+ voices — still after lock.

    Does this lip-sync each speaker?

    No. Audio only. Mouth work is a different tool on locked portraits.

    What if I only have the script and no picture yet?

    Write the script. Do not generate the voices. Picture first, or at least a duration contract. Voices second.

    Will a cloned house voice work as one of the speakers?

    If the TTS path you pick accepts that voice_id, yes — on the keeper. Cloning is setup. Casting is finishing.