Cast a multi-voice track against a locked cut, not a draft
generate_multi_speaker_speech builds a conversation in one call. Attaching it before picture lock orphans a timed dialogue on a recut.
A two-hander is one audio file, not two glued VO jobs. generate_multi_speaker_speech assigns a distinct voice to each speaker alias in a single call. attach_audio_to_video lays the result on a clip. That conversation has internal timing. Recut the picture and the Guest answers a question the Host no longer asks on screen. You will regenerate the whole cast.
Add multi-speaker dialogue to a video is the job. Billing follows script length and speaker count. estimate_cost before a long conversation. A single narrator is still generate_speech. Reach for the cast only when you actually need two or more voices.
The timing is the asset
Each alias in the script needs a matching entry in speakers. Miss that and the assignment is ambiguous. Get it right and you still have a track whose pauses assume a picture. Podcast-style clips, interview ads, scripted banter — all of them are spotted against this cut. A draft sequence with placeholder B-roll is not a cut.
People generate the dialogue first because the script is fun. Then they generate picture to "cover" it. Then the picture is 20% shorter. Then they regenerate the conversation. The first generate was a table read you paid model rates for. Table-read in a doc. Pay the tool when the cut can hold the lines.
Attaching a multi-speaker track to a miss also burns the attach fee. The speech file might be reusable if duration matches a later keeper. Duration rarely matches. Plan to regenerate the conversation against the locked file, which means do not generate it yet.
Visuals are a separate later step
This tool returns audio. On-screen speakers, split-screen, lipsync per person — none of that is included. Driving two portraits with generate_lipsync against the same conversation is two more models, each of which needs locked stills and this locked track. Stacking those on a draft is how a "quick podcast ad" becomes four finishing passes on files you will replace.
The AI video generator still has to produce plates worth covering. Dialogue is not coverage.
Order
- Lock picture duration, or accept MOS plates with a duration you will not change.
- Generate the multi-speaker track with aliases that match
speakers. - Attach with replace or mix, on purpose.
- Caption the conversation, not the silent picture you had yesterday.
FAQ
Can I chain two generate_speech calls instead?
You can, and you will hand-stitch timing yourself. When the picture changes you restitch. One multi-speaker call is the right tool for 2+ voices — still after lock.
Does this lip-sync each speaker?
No. Audio only. Mouth work is a different tool on locked portraits.
What if I only have the script and no picture yet?
Write the script. Do not generate the voices. Picture first, or at least a duration contract. Voices second.
Will a cloned house voice work as one of the speakers?
If the TTS path you pick accepts that voice_id, yes — on the keeper. Cloning is setup. Casting is finishing.