Converting a recorded take to another voice
An intake check, an input-prep chain and a three-way decision for when voice conversion beats re-recording or generating the line from text.
You have a recording. The words are right, the pacing is right, and the voice is wrong. Three routes out of that, and the useful way to compare them is by what each one destroys.
Re-record destroys the schedule. It also destroys the take, because a second performer will not independently rediscover the pause the first one put before the punchline, and neither will the same performer at 6pm on a Friday.
Re-synthesise from text destroys the performance. Text-to-speech starts from words and invents a delivery, because plain text does not contain one. Whatever made the original read work — where the stress landed, how long the beat held — was never in the script, so it cannot come back out of it.
Convert destroys nothing except the timbre, which was the thing that was wrong. Speech-to-speech takes a real delivery and swaps only the vocal identity performing it.
The concept is covered in re-voicing a recording without re-recording. This post is the practical layer: how to tell whether a specific take will survive the swap, what to do to the file before you send it, and where conversion is quietly the wrong tool.
The intake check
Five questions. A take that fails any of them will convert and produce something you throw away.
- Is the delivery actually right? Conversion preserves the performance in the source, flaws included. A flat read converts into a flat read in a different voice. If the take was mistimed, the problem is not the voice and conversion is not the fix.
- Is the vocal isolated, or is there music and room tone under it? A conversion works on the signal it is given. A take with a music bed baked in carries that bed into the conversion, and the result is a re-voiced line sitting on a smeared version of the original mix.
- Is there one speaker? Conversion targets a voice. A take with two people talking, or with an overlap where one person laughs over the other, does not have a single voice to swap. Split it or accept a mess.
- Are the lips on camera? Conversion is an audio operation. It does not re-sync mouth movement, so a take where the speaker is on camera and clearly visible will end up with lips timed to a voice that is no longer there. Narration, voiceover and off-camera dialogue are safe; a close-up talking head is not.
- Do you have the right to put those words in that voice? The performance is one person's and the target identity is potentially another's. If the target is a cloned voice of a real person, the consent question attaches to the output, not just to the clone.
Four yeses and a clean consent answer means the take is convertible. Anything else needs fixing before conversion, not after.
Prep the input
The conversion tool takes an audio URL. Not a video file, and not a rough mix. Two prep steps do almost all of the quality work.
Get the dialogue out as audio. If the take currently lives inside a video, the audio has to reach the tool as an audio file. This is the step that surprises people on their first attempt, and it is worth knowing before you start rather than after a failed call.
Isolate the vocal if anything is sitting under it. Vocal isolation strips background music and instrumentation from a mixed recording and hands back the voice on its own. Run it before conversion, not after — converting first and isolating second means the isolation is now working on a synthesised voice that was already contaminated by the bed. Isolating vocals from an audio clip is the direct route, and the voice isolation glossary entry covers what the operation does and does not recover. If the underlying take is noisy rather than merely mixed, rescuing noisy voiceovers with audio isolation is the longer treatment.
Two smaller habits that pay off:
- Trim to the line, not the session. Convert the segment you actually need. A three-minute file with twelve seconds of usable take in it converts all three minutes.
- Keep the original. Conversion is a generation, not an edit. The source file is your only way back if the target voice turns out wrong, and it is also the input for a second attempt at a different voice.
Picking the target voice
Two sources, and they behave differently in practice.
A catalog voice. Pick a voice ID from the Cartesia roster and convert into it. Fast, no setup, and the right choice when the target is an archetype rather than a person — a warmer narrator, a deeper register, a different age impression.
A cloned voice. A voice cloned from an audio sample returns a reusable ID that works both for generating speech from text and for converting an existing take. This is the chain that makes conversion strategically interesting rather than merely convenient: one cloned identity can speak new scripts and absorb existing performances. Cloning a voice from a recording is where that ID comes from.
Your cloned-voice listing returns newest first and caps at 30. For a working set that is plenty. For a team that clones a voice per project and never prunes, it is a reason to name them properly and retire the finished ones, because a target voice you cannot find is a target voice you do not have.
There is also a second engine for this job. The agent's conversion tool runs on Cartesia; Eleven Labs Voice Change sits in the model catalog as its own voice-conversion model. Neither is a strictly better default — they pull from different voice libraries, and which one fits depends on where the voice you want actually lives.
Run it, then put it back
The conversion itself is one call with three inputs: the audio URL, the target voice ID, and an output format. Everything interesting already happened in prep.
Putting the result back into a video is a separate step with a decision in it. Attaching audio to a video has two modes and picking the wrong one is the most common post-conversion mistake:
- Replace strips the video's original audio and substitutes the new track. This is the standard case for a straight voice swap where the converted line should be the only thing audible.
- Mix keeps the original audio and layers the new track over it, with independent volume controls for each. This is the mode when music or sound effects on the original track need to survive underneath the re-voiced line.
Choose replace when the original audio was only ever the take. Choose mix when the original track carries anything you would have to rebuild otherwise. The video stream itself is copied without re-encoding, so the visuals do not degrade in the process, and the operation carries a flat credit fee regardless of clip length. Changing the voice in a video walks the same chain as a task, and replacing video audio covers the attach step on its own.
One verification before you call it done: transcribe the converted audio and read it. A transcript is a fast, objective check that nothing got mangled where the source was quiet or overlapping, and it catches the failure you will listen straight past on the third playthrough.
Four takes that should not be converted
- A close-up talking head. The lips are the problem. This is a lip-sync job or a re-shoot, not an audio swap.
- A take where the performance is the complaint. If someone said "it sounds flat," they are not describing timbre. Conversion will hand you the same flat read in a new voice and the note will come back unchanged.
- A two-hander with overlap. One target voice cannot resolve two speakers. Split the takes if they are separable and treat them as two jobs.
- A line that has not been approved yet. Conversion is a finishing operation. Converting a take whose copy is still in review means converting it again after the review, which is a second generation for a line you were always going to change.
Conversion is narrow and very good inside its lane. It is the only one of the three routes that treats the existing performance as an asset rather than something to replace, and that is worth exactly as much as the performance is. On a great take, a lot. On a bad one, nothing, and the tool will not tell you which you have.
FAQ
Do I need to isolate the vocal if the recording is already clean?
No. Isolation is for takes with music, room tone or instrumentation sitting under the voice. A clean, close-mic'd recording with nothing behind it goes straight to conversion. Running isolation on an already-clean file is an extra generation for no gain, and on a very clean source it can take something out that you wanted to keep.
Will the converted voice keep the original accent?
Conversion swaps vocal identity while preserving the delivery, so the timing, stress and phrasing of the source carry through. The target voice's own characteristics come with it. If the source performance and the target voice pull in different directions on accent, the result can sit somewhere between the two — worth testing on one line before committing a whole session.
Can I convert a take into a language it was not recorded in?
No. Conversion changes who is speaking, not what is said. Translating content is dubbing, which is a different operation with its own timing implications because the translated line will not be the same length as the original.
How do I choose between conversion and cloning the original speaker?
They answer different questions. Conversion changes the voice on one existing take. Cloning captures a voice identity so it can speak scripts that do not exist yet. If you need this one file in a different voice, convert. If you need a voice you will use repeatedly across future content, clone — and note that a cloned ID then works as a conversion target too, so the two are not mutually exclusive.