Two levels exist in practice. A quick clone works from a short sample — often under a minute — and captures timbre convincingly while being looser on the speaker's habits of rhythm and emphasis. A high-fidelity clone wants substantially more audio and gets closer to a specific person's delivery, not just their tone.
Sample quality dominates everything else. One speaker, one room, no music, no overlap, consistent distance from the microphone, and enough range that the model hears more than one emotional register. Ten clean minutes beat an hour of podcast audio with a co-host in it, because the model treats whatever is consistently present as part of the voice.
The obvious caveat is not a technicality: a clone of someone else's voice needs their permission, and both the platform's terms and the law in most jurisdictions treat it that way. Recorded consent from the speaker is the norm for commercial use.
In practice
- One speaker, clean room, no background music — everything else compromises the clone.
- Include varied delivery in the sample: the model can only reproduce registers it heard.
- A clone is a reusable asset — build it once carefully rather than repeatedly in a hurry.
Voice cloning models
Catalog entries that build a reusable voice from a sample you provide. 5 of the 296 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Cartesia Sonic 3.5 | Cartesia | Audio |
| Qwen 3 TTS Voice Design | Qwen | Audio |
| Cartesia Voice Clone | Cartesia | Audio |
The mistake to avoid
Cloning from a video's existing soundtrack. Ambience, music and room reverb are learned as part of the voice and turn up in every line it later speaks.
Where you will run into it
- Clone Your Voice for Videos — Record once. Reuse the voice forever.
- AI Voice Cloning & Text to Speech — Your voice. Any language. Any script.
Related terms
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
Speech-to-speech
Speech-to-speech takes a recording of one person talking and re-renders it in a different voice, keeping the original performance intact.
Voice design
Voice design creates a new synthetic voice from a written description — age, accent, texture, energy — instead of cloning one from a recording.
AI dubbing
AI dubbing replaces a video's spoken audio with another language, usually keeping the original speaker's voice and optionally re-syncing their mouth.
Audio tags
Audio tags are markers written inside the text of a script — bracketed or angle-bracketed cues like a laugh or a whisper — that tell a speech model how to deliver the words around them.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.