Re-voice a performance you are keeping, not a take you will throw out
change_voice keeps timing and delivery, swaps the Cartesia speaker. Re-voicing a miss, or re-voicing before picture lock, pays for a performance on frames you will replace.
change_voice is for when the performance is right and the speaker is wrong. It converts an existing audio clip to a different Cartesia voice — including a clone — without letting you type new words. Timing stays. Delivery stays. Then attach_audio_to_video puts the new speaker back on the footage. If the footage is a miss, you paid to recast a take you will delete. If you still need different words, this is the wrong tool entirely.
Change the voice in a video is the job. Typical flow: get the dialogue as standalone audio, run change_voice, reattach. transcribe_audio is how you inspect the words you are about to keep. Billing is per clip through change_voice plus the flat attach. estimate_cost for the pair.
Words stay. Picture must stay too.
This is not a rewrite. If the offer changed, generate new speech. If the take is getting regenerated because the eyes are wrong, wait. The new take will have new timing even if you plan to replace its audio; attaching a recast of the old performance to a new picture is a lip-flap lottery unless the shot is MOS B-roll.
People recast early because the temp voice "isn't on brand." Brand is a later problem than "is this the shot." Recasting a candidate also trains the team to treat change_voice as a preview button. It is a process on a file.
Cartesia targets only. Your cloned voice_id is in scope. Other families are not.
Order
- Lock the picture, or confirm the shot is MOS and duration-stable.
- Confirm the recorded words are the words you will ship. Transcribe if you need to see them.
change_voiceon that audio.- Attach. Then caption the new speaker, not the old one.
If you needed a different line at 0:06, you needed an edit or a new read, not a recast.
The voice cloning path is how you get a target speaker. Recast is how you apply that speaker to a keeper performance.
Do not recast as a way to A/B voices on a maybe
Three Cartesia targets on one unlocked take is three bills and no decision about the picture. A/B voices on a locked MOS product shot is a real test. Sequence matters.
FAQ
Can I type a correction while changing voice?
No. Same words, new speaker. Corrections are generate_speech with a new script, after you know the cut.
Does this run on the video file directly?
It runs on audio. You pull or isolate the track, recast, reattach. Plan that pipeline on a keeper so you are not isolating a miss.
Should I isolate noisy VO before recasting?
If the keeper's mix is dirty, yes — isolate_audio on the file you will ship, then recast. Isolating a throwaway to recast it is two finishing tools on a ghost.
Why not just generate_speech in the new voice?
Because you would lose the original timing and delivery. change_voice exists when those are the thing you liked. That liking should be about a take you are keeping.