clone_voice_from_audio takes a sample recording and returns a voice_id — a reusable Cartesia voice you can call on for every future generate_speech or change_voice job, instead of picking from the stock catalog or re-recording each time.
This is the foundation for scaling narration in your own voice across many videos without re-recording, and for keeping voice identity consistent across a series.
Powered by
Clone a reusable Cartesia voice from an audio sample. Returns a voice_id usable with generate_speech (Cartesia models) and change_voice. Use when the user wants to create a voice from a recording rather than pick from the existing catalog.
What to tell the agent
Versely's agent maps this plain-English request directly onto clone_voice_from_audio. You don't need to know the parameter names — just describe what you want.
“Clone my voice from this audio sample and name it 'Main Narrator', then use it to read this script.”
How it works
1. Provide a clean sample
Upload or point to an audio_url with clear, representative speech — the cleaner the sample, the better the clone.
2. Name and describe it
name is required; description and language are optional but help you tell voices apart later.
3. Get your voice_id back
The returned voice_id is what you pass to generate_speech (Cartesia models) for narration, or change_voice to re-voice existing recordings.
4. Reuse across projects
Browse your saved voices anytime with list_cloned_voices instead of re-cloning for every new video.
What it costs
Cloning is billed once per sample submitted; reusing the resulting voice_id in generate_speech is billed at that tool's normal per-generation rate. Confirm exact pricing with estimate_cost.
The formula behind that number: What does an AI voiceover cost for my script? — meters the characters you type, not the seconds you get back.
Limits & things to know
- Quality depends heavily on the sample's audio cleanliness — background noise or multiple speakers in the sample will degrade the clone.
- The cloned voice is only usable with Cartesia-backed calls (generate_speech with a Cartesia model, and change_voice) — it isn't a universal voice_id across every TTS model in the app.
- Only clone voices you own or have explicit consent to use.
Who uses this
- Narrating a whole video series in your own voice without re-recording
- Personal-brand YouTube channels
- Consistent voice identity across a UGC ad campaign
- Preserving a voice for future dubbing or narration work
Frequently asked questions
How much audio do I need to clone a voice?+
A clean, representative sample is enough — clone_voice_from_audio doesn't require a long recording, but clarity matters more than length.
Where can I use the cloned voice afterward?+
The returned voice_id works with generate_speech (on Cartesia models) for new narration, and with change_voice to re-voice an existing recording.
Can I see all my cloned voices later?+
Yes — list_cloned_voices returns your saved voices so you can reuse one without cloning again.
Can I clone someone else's voice?+
Only clone voices you own or have explicit consent to use — third-party or celebrity voices require verification and disallowed use cases are blocked.
Related Versely tools
Related editing jobs
Clone Your Voice for Videos inside Versely
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.