Audio, Voice & Dubbing · Versely AI

    Add Multi-Speaker Dialogue to a Video

    A whole cast, one API call.

    Interviews, podcast-style clips, and scripted conversations need more than one voice — and chaining several single-voice generate_speech calls means manually stitching the timing together yourself. generate_multi_speaker_speech handles this in one call: each speaker alias gets its own distinct voice, spoken in the order their lines appear.

    Once the dialogue track is generated, attach it to your video with attach_audio_to_video the same way you would a single voiceover.

    Powered by

    generate_multi_speaker_speechattach_audio_to_video

    Generate multi-voice dialogue/podcast audio in one call (Gemini multi-speaker TTS) — each speaker alias gets its own distinct voice, spoken in the order their lines appear in `text`. Use for scripted conversations, interviews, or podcast-style clips with 2+ distinct voices instead of chaining multiple generate_speech calls.

    What to tell the agent

    Versely's agent maps this plain-English request directly onto generate_multi_speaker_speech. You don't need to know the parameter names — just describe what you want.

    Generate a two-speaker dialogue for this script with distinct voices for Host and Guest, then attach it to the video.

    How it works

    1. 1. Script the conversation

      Write each line with a speaker alias attached — e.g. 'Host:' and 'Guest:' — so the order and assignment are unambiguous.

    2. 2. Define the speakers

      The speakers parameter assigns a distinct voice to each alias used in the script.

    3. 3. Set language and tone

      language and style_prompt steer the overall delivery — pacing, formality, energy.

    4. 4. Attach to video

      Feed the resulting audio into attach_audio_to_video (mode 'replace' or 'mix' depending on whether the video already has usable sound).

    What it costs

    Billed per generation based on script length and speaker count — ask the agent for estimate_cost before generating a long conversation.

    The formula behind that number: What does an AI voiceover cost for my script? meters the characters you type, not the seconds you get back.

    Limits & things to know

    • Each speaker alias in your script needs a matching entry in the speakers parameter, or the assignment will be ambiguous.
    • This generates the audio only — pairing it with on-screen visuals or lipsync per speaker is a separate step.
    • For a single narrator, generate_speech is the simpler tool — reach for this one specifically when you need 2+ distinct voices.

    Who uses this

    • Podcast-style video clips
    • Scripted interview formats
    • Two-character explainer dialogues
    • Debate or panel-style content

    Frequently asked questions

    Do I need to generate each voice separately?+

    No — generate_multi_speaker_speech produces the full conversation in one call, with each speaker alias getting its own distinct voice in the order their lines appear in the script.

    How many speakers can it handle?+

    You define one voice per speaker alias via the speakers parameter — the tool is built for scripted conversations with 2+ distinct voices.

    How do I get this onto my video?+

    Generate the dialogue audio first, then attach it with attach_audio_to_video the same way you would a single-voice narration.

    See it in a workflow

    Related Versely tools

    Related editing jobs

    Add Multi-Speaker Dialogue to a Video inside Versely

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.