What you say to the agent
No special syntax — just describe it like you would to a person.
What it does, step by step
- 1
Describe the voice character in words instead of picking from a fixed voice_id list.
- 2
Give it the text you want spoken in that voice.
- 3
It designs the voice (Qwen 3 TTS Voice Design) and speaks your text with it in one call.
What it needs from you
- •The text/script you want spoken
- •A text prompt describing what you want
What comes back
One audio clip in the described voice.
What it costs
Priced per generation, shown before you confirm.
Good to know
- This doesn't create a reusable voice_id for other tools — re-describe the same voice to get a similar sound on new text. For a reusable voice, use voice cloning from an audio sample instead.
Under the hood
This is what the agent actually calls when you ask for it — real tools from its live surface, not marketing copy.
create_custom_voiceDesign a custom voice from a text style prompt (Qwen 3 TTS Voice Design) and speak the given text with it in one call. Unlike generate_speech's fixed voice_id list, the voice itself is described in words (e.g. 'a gravelly older narrator with a slow, confident cadence'). Returns one audio clip synchronously — this does not create a reusable voice_id for other tools; re-run with the same prompt for a similar-sounding voice on new text.
The full tool behind it
See it done in a real workflow
Morning Matcha Routine
UGC-style wellness reel — a clean-girl creator walks through her actual morning matcha ritual, then cuts to a cozy animated hero shot of the finished iced latte.
TrendingNYC Street Interview
Photorealistic street-vlog interview — Riley works three different NYC corners at golden hour asking strangers one question: "What's the wildest thing you've ever done?" Three candid OTS/two-shot clips with locked character references, real handheld energy and native spoken dialogue. Vertical 9:16.
TrendingPrimo Protein vs Other Brand
Pixar-style 3D comparison ad — the confident PRIMO PROTEIN pouch faces off against the tired old rival brand's tub across 11 talking clips: pasture vs dusty pantry, herb garden vs toxic lab, clean American lab vs grimy factory, chocolate-milkshake CTA vs "wet sand." ~40s vertical reel with native character voices.
Or start from a one-tap template
Frequently asked questions
What do I actually say to the agent to design a custom AI voice from a description?+
Just describe it in plain English — for example: "Design a gravelly older narrator with a slow, confident cadence and read this intro" The agent handles picking the right tool and model from there.
What does the agent need from me first?+
At minimum: The text/script you want spoken; A text prompt describing what you want. Anything else it needs, it asks for before running.
What do I get back?+
One audio clip in the described voice.
Does this cost credits?+
Priced per generation, shown before you confirm.
Anything I should know before asking for this?+
This doesn't create a reusable voice_id for other tools — re-describe the same voice to get a similar sound on new text. For a reusable voice, use voice cloning from an audio sample instead.
You can also just ask for
Ask your Versely agent to design a custom AI voice from a description
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.