There is no source speaker. You describe the voice you want in the same way you would brief a casting director, and the model synthesises something matching it. That makes the output nobody's voice in particular, which sidesteps the consent question that cloning necessarily raises.
It is the right tool for characters and brand voices that should not be traceable to a person: a narrator for a faceless channel, a mascot, a set of distinct voices for an animated cast. It is the wrong tool when the point is that a specific real person is speaking.
Descriptions behave like prompts, with the same quirks. Concrete, physical attributes — pitch, pace, warmth, breathiness, accent — steer reliably; abstract adjectives like "trustworthy" mostly do not. And a designed voice is only as reusable as your ability to reproduce it, so save the description alongside the result.
In practice
- Describe physical qualities, not personality traits.
- Save the exact description — it is the only record of how the voice was made.
- Generate a few candidates and audition; description-to-voice is not one-to-one.
The mistake to avoid
Designing a voice per project. Distinct channels need distinct voices, but one channel needs one voice used consistently for months.
Where you will run into it
- AI Voice Cloning & Text to Speech — Your voice. Any language. Any script.
Related terms
Voice cloning
Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
Voice stability
Voice stability is the control that decides how much a synthetic voice varies its delivery — steady and predictable at one end, expressive and unpredictable at the other.
Audio tags
Audio tags are markers written inside the text of a script — bracketed or angle-bracketed cues like a laugh or a whisper — that tell a speech model how to deliver the words around them.
Speech-to-speech
Speech-to-speech takes a recording of one person talking and re-renders it in a different voice, keeping the original performance intact.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.