What you say to the agent
No special syntax — just describe it like you would to a person.
What it does, step by step
- 1
Write your script with each line attributed to a speaker alias.
- 2
Give the agent the text plus the list of speakers — each alias gets its own distinct voice, spoken in order.
- 3
It generates one audio file with the full conversation, saving you from chaining multiple single-voice calls.
What it needs from you
- •The text/script you want spoken
- •Who's speaking and which lines belong to whom
What comes back
A single audio clip containing the full multi-speaker conversation.
What it costs
Priced per generation, shown before you confirm.
Under the hood
This is what the agent actually calls when you ask for it — real tools from its live surface, not marketing copy.
generate_multi_speaker_speechGenerate multi-voice dialogue/podcast audio in one call (Gemini multi-speaker TTS) — each speaker alias gets its own distinct voice, spoken in the order their lines appear in `text`. Use for scripted conversations, interviews, or podcast-style clips with 2+ distinct voices instead of chaining multiple generate_speech calls.
The full tool behind it
See it done in a real workflow
Podcast Clip — Versely
Podcast-clip style talking-head reel — a creator in a cozy studio teases her income jump, then reveals Versely as the tool that lets her ship videos without an editor.
PodcastPodcast with me
Host invites guests on his podcast and talks about the future and Artificial Intelligence.
TrendingMorning Matcha Routine
UGC-style wellness reel — a clean-girl creator walks through her actual morning matcha ritual, then cuts to a cozy animated hero shot of the finished iced latte.
Or start from a one-tap template
Frequently asked questions
What do I actually say to the agent to create a multi-voice dialogue or podcast clip?+
Just describe it in plain English — for example: "Generate a podcast intro with two hosts, Alex and Jamie, riffing about AI video tools" The agent handles picking the right tool and model from there.
What does the agent need from me first?+
At minimum: The text/script you want spoken; Who's speaking and which lines belong to whom. Anything else it needs, it asks for before running.
What do I get back?+
A single audio clip containing the full multi-speaker conversation.
Does this cost credits?+
Priced per generation, shown before you confirm.
You can also just ask for
Ask your Versely agent to create a multi-voice dialogue or podcast clip
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.