change_voice re-voices a take; it does not write a new read
Voice change is one named audio job: same delivery, different Cartesia voice. TTS from scratch, cloning, and dubbing are other bills.
change_voice re-voices a take. Same delivery, same timing, different Cartesia voice. It does not write a new read from a script. It does not clone a reusable voice_id from a sample. It does not dub a video into French. If you re-typed the line into TTS because you wanted a deeper voice, you spent a new performance and threw away the one you had already approved.
Change the voice in an audio clip is the named agent job. Required: audio_url and voice_id. Optional output format. The tool converts the voice while preserving the original delivery — a Cartesia voice-change pass — rather than re-speaking the line from scratch. Priced per generation, shown before you confirm.
The AI voice cloning surface is how you obtain a voice_id worth pointing at. This capability is the swap, not the clone, not the first read.
What the job is
An existing recording you like, except the timbre is wrong. A cloned voice you already saved. Or a catalog Cartesia voice. Output is the same audio, re-voiced. Attach it to video later with add music or a voiceover if the picture is already right.
If you need the clone to exist first, clone my voice from a recording returns a reusable voice_id for generate_speech and for change_voice. Clone once. Swap many times. Do not clone on every swap.
What it is not
- Text-to-speech. Write and generate a voiceover is
generate_speech. New words, new timing, new performance. Use it when the copy changed. Usechange_voicewhen the copy did not. - Multi-speaker dialogue. Create a multi-voice dialogue generates the conversation in one take. Re-voicing one mixed two-hander as if it were one person will smear both aliases.
- Dubbing. Dub a video is language plus (optionally) lipsync, off hosted media, async. Voice-change does not translate.
- Vocal isolation. Pulling a voice out of a mix is
isolate_audio. Change needs a voice it can convert, not a song with a singer somewhere in it.
The waste pattern is almost always TTS-from-scratch on a take whose pacing you already fought for. You will not get that pacing back for free.
The test
Do you already have the read, and is the only problem who it sounds like?
Yes: this job. Pick the voice_id. Confirm.
If the words are wrong, regenerate speech. If the person must be seen, lipsync after the new audio. If the language is wrong, dub. Name those jobs; do not stretch this one.
FAQ
Can I change the voice and rewrite a sentence in the same pass?
Rewriting is a new read. Change will keep the old timing, including the sentence you no longer want. Generate speech for the new copy.
Does this work on video, or only on audio files?
The tool takes audio_url. Strip or export the voice, convert, then attach. Do not send a full muxed social MP4 and expect a clean swap.
Is a cloned voice required?
No, but you need some Cartesia voice_id. Cloning is how you get your id. Catalog ids work for a different person.
Why not always regenerate TTS in the new voice?
Because you lose the take. Breath, pace, the laugh you kept. Voice-change exists for the case where those are the asset.