Cloning a voice from a recording is a reusable voice, not a one-off read
clone_voice_from_audio mints a voice_id you keep using. It is not a TTS pass and not a voice you describe in words.
Cloning a voice from a recording is a reusable voice, not a one-off read. One recording. A reusable voice from then on. clone_voice_from_audio takes a sample, a name, and returns a Cartesia voice_id you can call later for voiceover or a voice change. It does not speak your script. Using a clone job as a TTS pass — or describing a voice in adjectives and calling that a clone — wastes credits on the wrong audio row.
Clone my voice from a recording is the named capability. The AI voice cloning tool is the same family on the studio side. list_cloned_voices is how you find the name you already minted, capped at a list you should not clutter with retries of a bad file.
What you get back
You attach an audio sample of the voice. You name it something you can find in a month (Weekly Recap Host, not clone 3). The agent clones a reusable Cartesia voice and returns a voice_id. From then on you reference that name in a script-to-voiceover request or a voice change. You do not re-upload the sample every Tuesday.
Priced per clone, shown before you confirm. The charge is for the voice object, not for a line of copy. If you needed one read of one paragraph, you wanted speech generation, not a clone.
What wastes credits
No sample, just adjectives. "Give me a gravelly older narrator" is design a custom AI voice. That path designs from words and speaks the line in one call. It does not mint a reusable voice_id. Re-describe it next time if you want it again. A clone is the opposite contract: a recording in, an ID out, no script required.
You wanted the words spoken now. Paste the script into voiceover. Cloning first, then immediately generating one line you could have spoken from a catalog voice, is two charges to skip picking a preset.
A noisy, mixed, or too-short sample. Cloning copies what you give it. A voice memo in a cafe with a song underneath is how the clone sounds like a cafe. Isolate or re-record before you mint. list_cloned_voices fills up with retries of the same bad file.
Swapping a voice on an existing take. If the performance is already good and you only want a different throat on it, that is change_voice, not a new clone.
The test
Do you have a recording you would actually reuse across videos?
If yes, clone it once, name it, stop. Future scripts point at the name.
If you have a script and no recording, generate speech. If you have a recording and you only need this one file re-voiced, change the voice. If you have adjectives and no audio, design a voice for that line. This is one named agent job. Using it for a different job wastes credits.
FAQ
Do I lose the clone if I generate the next voiceover in a new chat?
No. The voice_id is reusable. Refer to it by the name you gave it. The agent can list cloned voices when you forget the label.
Is cloning the same as reading my script in "my" voice without a sample?
No. Without a recording there is nothing to clone. A described voice is a one-shot design. A clone is a sample-backed ID.
Can I clone someone else's voice from a podcast download?
Only with rights you actually have. The tool will mint from the file you attach. Clearance is your problem, not the model's.
Why not clone a new voice for every video?
Because the job is reuse. A new clone per episode is how you get twelve cousins of the same throat and a cluttered voice list. Mint one keeper. Drive it with new scripts.