The script carries speaker aliases — Host, Guest — and the speakers list assigns a voice to each alias. One call emits one file in line order. Chaining single-voice reads means you own the gaps, the overlaps and the stitch; this call does not.
Text-to-speech is one voice reading a script, and that page already points here when the brief is a cast. Lipsync drives a mouth; it does not cast two voices. Audio-to-video conditions picture on a soundtrack you already have. Voice design can invent the voices you then assign to aliases, which is casting, not the conversation generate.
Attach the file to picture afterwards. Pairing each speaker with on-screen lipsync is a later step, not part of the dialogue generate.
In practice
- Attribute every line to an alias that also appears in the speakers list — an unmatched name is an ambiguous assignment.
- Write the conversation as it should be heard, including overlap you actually want; the model will not invent a cross-talk you did not script.
- Use the single-voice job for a narrator; reach for this one only when two or more distinct voices are the point.
The mistake to avoid
Generating each role as a solo read and stitching. Timing and overlap were the whole reason for one call, and two monologues still sound like two monologues.
Go deeper
generate_multi_speaker_speech is one conversation, not stacked solos
Multi-speaker TTS is one named audio job. Chaining generate_speech and calling it a podcast wastes takes and still sounds like two monologues.
Where you will run into it
- Add Multi-Speaker Dialogue to a Video — A whole cast, one API call.
- AI Voice Cloning & Text to Speech — Your voice. Any language. Any script.
Related terms
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
Lipsync
Lipsync generation drives a face's mouth from an audio track, so the speech reads as spoken rather than dubbed over the top.
Audio-to-video
Audio-to-video generation drives the picture from a soundtrack: the audio is the primary input and the visuals are generated to agree with it.
Voice design
Voice design creates a new synthetic voice from a written description — age, accent, texture, energy — instead of cloning one from a recording.
Voice cloning
Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.