What you say to the agent
No special syntax — just describe it like you would to a person.
What it does, step by step
- 1
Describe the voice character in words instead of picking from a fixed voice_id list.
- 2
Give it the text you want spoken in that voice.
- 3
It designs the voice (Qwen 3 TTS Voice Design) and speaks your text with it in one call.
What it needs from you
- •The text/script you want spoken
- •A text prompt describing what you want
What comes back
One audio clip in the described voice.
What it costs
Priced per generation, shown before you confirm.
Good to know
- This doesn't create a reusable voice_id for other tools — re-describe the same voice to get a similar sound on new text. For a reusable voice, use voice cloning from an audio sample instead.
Under the hood
This is what the agent actually calls when you ask for it — real tools from its live surface, not marketing copy.
create_custom_voiceDesign a custom voice from a text style prompt (Qwen 3 TTS Voice Design) and speak the given text with it in one call. Unlike generate_speech's fixed voice_id list, the voice itself is described in words (e.g. 'a gravelly older narrator with a slow, confident cadence'). Returns one audio clip synchronously — this does not create a reusable voice_id for other tools; re-run with the same prompt for a similar-sounding voice on new text.
The full tool behind it
See it done in a real workflow
Morning Matcha Routine
UGC-style wellness reel — a clean-girl creator walks through her actual morning matcha ritual, then cuts to a cozy animated hero shot of the finished iced latte.
TrendingNYC Street Interview
Photorealistic street-vlog interview — Riley works three different NYC corners at golden hour asking strangers one question: "What's the wildest thing you've ever done?" Three candid OTS/two-shot clips with locked character references, real handheld energy and native spoken dialogue. Vertical 9:16.
TrendingPrimo Protein vs Other Brand
Pixar-style 3D comparison ad — the confident PRIMO PROTEIN pouch faces off against the tired old rival brand's tub across 11 talking clips: pasture vs dusty pantry, herb garden vs toxic lab, clean American lab vs grimy factory, chocolate-milkshake CTA vs "wet sand." ~40s vertical reel with native character voices.
Or start from a one-tap template
AI Inflate
Puff up like a balloon in the middle of a park, lift off your feet and float away into the sky.
Earth Zoom Out
Start on you looking up at the camera, then fly straight up past the rooftops and clouds until the whole Earth is in frame.
Future Baby
Upload two photos and see what your baby could look like, in a warm family moment at home.
Frequently asked questions
What do I actually say to the agent to design a custom AI voice from a description?+
Just describe it in plain English — for example: "Design a gravelly older narrator with a slow, confident cadence and read this intro" The agent handles picking the right tool and model from there.
What does the agent need from me first?+
At minimum: The text/script you want spoken; A text prompt describing what you want. Anything else it needs, it asks for before running.
What do I get back?+
One audio clip in the described voice.
Does this cost credits?+
Priced per generation, shown before you confirm.
Anything I should know before asking for this?+
This doesn't create a reusable voice_id for other tools — re-describe the same voice to get a similar sound on new text. For a reusable voice, use voice cloning from an audio sample instead.
You can also just ask for
Ask your Versely agent to design a custom AI voice from a description
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free account. Works in your browser - no install needed. The same account signs in on your phone.