Speech, voice and audio

    Multi-speaker dialogue

    Also called Multi-voice dialogue, Multi-speaker TTS.

    Multi-speaker dialogue is one speech generation with a distinct voice per speaker alias, so a conversation is a single take instead of several stitched reads.

    The script carries speaker aliases — Host, Guest — and the speakers list assigns a voice to each alias. One call emits one file in line order. Chaining single-voice reads means you own the gaps, the overlaps and the stitch; this call does not.

    Text-to-speech is one voice reading a script, and that page already points here when the brief is a cast. Lipsync drives a mouth; it does not cast two voices. Audio-to-video conditions picture on a soundtrack you already have. Voice design can invent the voices you then assign to aliases, which is casting, not the conversation generate.

    Attach the file to picture afterwards. Pairing each speaker with on-screen lipsync is a later step, not part of the dialogue generate.

    In practice

    • Attribute every line to an alias that also appears in the speakers list — an unmatched name is an ambiguous assignment.
    • Write the conversation as it should be heard, including overlap you actually want; the model will not invent a cross-talk you did not script.
    • Use the single-voice job for a narrator; reach for this one only when two or more distinct voices are the point.

    The mistake to avoid

    Generating each role as a solo read and stitching. Timing and overlap were the whole reason for one call, and two monologues still sound like two monologues.

    Go deeper

    generate_multi_speaker_speech is one conversation, not stacked solos

    Multi-speaker TTS is one named audio job. Chaining generate_speech and calling it a podcast wastes takes and still sounds like two monologues.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.