Voice Changing: Re-Voice a Recording Without Re-Recording
A guide read has perfect timing and the wrong voice. Re-voicing keeps the performance and swaps only the timbre — retyping the line throws both away.
A guide vocal comes back from a scratch session and it's exactly right — the pause before the punchline lands, the slight rasp on the emphasis word, the pacing that makes the whole read feel unforced instead of narrated. It's just in the wrong voice. The usual fix is to throw the take away, type the line out as text, and generate a fresh voiceover in the voice you actually wanted — which solves the voice problem and re-breaks the performance problem you'd already solved. Re-voicing is the tool for exactly this situation: it keeps the take and changes only the one thing that was actually wrong with it.
What gets thrown away when you retype the line instead
Text-to-speech and re-voicing sound like they're solving the same problem, and they're doing something structurally different. Handed a script, a text-to-speech model invents a delivery from nothing — it has to guess where the emphasis goes, how long to hold a pause, whether a line reads as sincere or sarcastic, because none of that exists anywhere in the plain text you gave it. Speech-to-speech works from the opposite direction: it's handed an actual delivery — a real recording, with real timing already baked in — and swaps only the voice identity, leaving the performance exactly as it was.
That distinction is the entire case for reaching for one over the other. If what you have is a script and no delivery yet, text-to-speech is doing necessary work — someone has to invent the performance, and that's its job. If what you already have is a take with the delivery you want, running it back through a text-to-speech generator throws that performance away and gambles on a fresh one landing just as well. It usually doesn't land the same, because the new generation has no idea what made the original take work in the first place.
What actually survives the swap
The practical claim worth being precise about: timing, emphasis, pacing, the pause held a beat too long on purpose, a laugh landing mid-sentence — all of it carries through, because none of it is what's being changed. Only the timbre, the specific vocal identity reading the line, gets replaced. It's closer to swapping an actor's voice in post while keeping every choice they made in the take than it is to hiring a new actor and hoping they make the same choices.
That's also the honest limit of the tool. It re-voices what's already there — it doesn't fix a flat read, it doesn't add energy that wasn't in the original performance, and it can't rescue a take where the underlying delivery was actually the problem. Re-voicing is for the case where the performance was right and the voice was wrong, not a general-purpose quality fix for a take that needed a better performance in the first place.
Three related tools, three different starting points
It's worth being clear about where re-voicing sits next to two tools it's easy to confuse it with, because each one assumes a different starting material:
- Text-to-speech starts from a script with no performance yet. It's the right tool when nothing has been recorded — the delivery gets invented from the text and whatever style guidance you provide.
- Voice cloning starts from a sample of someone's voice and builds a reusable identity from it, so future arbitrary text — scripts that don't exist yet — can be spoken in that voice later. It's an investment in a voice you'll reuse repeatedly, not a fix for one specific existing take.
- Re-voicing (speech-to-speech) starts from a specific recording you already have and already like the performance of. It's the narrowest of the three, and it's the right one exactly when the other two would be solving a problem you don't actually have — you don't need a script written from nothing, and you're not building a reusable voice identity, you just need this one take in a different voice.
Where this actually comes up
A few situations where a re-voice pass is the obvious tool rather than a workaround:
- A scratch or guide vocal recorded by a writer or director to lock in cadence and comedic timing, meant from the start to be swapped for the actual talent's voice once the read is approved — re-voicing is the entire point of a guide take, not an afterthought.
- A recording of someone no longer available — a past employee, a one-time guest — where new content needs the same distinctive cadence a cloned or reused voice can carry, applied to a fresh recording someone else performed with that cadence in mind.
- A UGC-style ad read where the writer performs the exact pacing an ad read needs — punchy, specific pauses before the offer — and that performance gets swapped onto the creator or actor's voice afterward rather than trusting a second performer to independently rediscover the same rhythm.
Two engines, one job
Versely has two separate paths to voice conversion, worth knowing apart. The agent's change_voice tool runs on Cartesia specifically — it converts an existing clip's voice to a different Cartesia voice while holding delivery and timing fixed, and it's the default path when you're working conversationally with the agent. Separately, Eleven Labs Voice Change sits in the model catalog as its own dedicated voice-conversion model, transforming voice characteristics while preserving speech content through a different provider's engine. Neither is a strictly better default — they're two different engines aimed at the same underlying job, and which one fits depends on which voice library you actually want the output voice pulled from.
A Versely walkthrough: re-voicing a line inside a finished video
The one wrinkle worth knowing in advance: change_voice operates on audio, not on a video file directly. Re-voicing dialogue that's already inside a finished video isn't a single step — it's a short chain, and knowing the chain up front saves a confused first attempt.
- Get the dialogue as a standalone audio clip. Start from the video's own audio track rather than the video file itself —
change_voiceneeds anaudio_url, not avideo_url. - Run the conversion. "Change this clip's voice to [target voice], keep the exact same timing and delivery" calls
change_voicewith your chosenvoice_id— from Cartesia's catalog directly, or a voice you cloned yourself beforehand. The pacing, pauses, and emphasis from the original take carry through untouched; only the voice identity changes. - Reattach the re-voiced audio to the video.
attach_audio_to_videoputs the new track back onto the original footage. Use replace mode when the re-voiced line should be the only audio — the standard case for swapping a guide vocal for a final one — or mix mode if the original track's music or sound effects need to survive underneath the new voice rather than being replaced entirely. - Spot-check timing against the visuals, not just the audio in isolation. Re-voicing preserves the original clip's own timing perfectly, but if the video has lip movement synced to the old voice, a re-voice pass doesn't re-sync lips to the new timbre — it's a straight audio swap, so it's the right tool for narration, voiceover, and off-camera dialogue, and a different tool's job for on-camera lip-synced speech.
FAQ
Does re-voicing change how fast or slow the line sounds?
No — the entire point is that timing carries through unchanged. If the original take has a two-second pause before the punchline, the re-voiced version keeps that exact pause; only the vocal identity delivering the line is different.
Can I re-voice a line into a voice I've never used before?
Yes, as long as the target voice is available in the catalog you're pulling from — Cartesia's voice library for the agent's built-in tool, or a voice you've cloned yourself and saved for reuse.
What if the original take's performance wasn't actually good?
Re-voicing won't fix that — it preserves whatever delivery is in the source recording, flaws included. A flat or mistimed original read needs either a better take or a text-to-speech generation with fresh style guidance, not a voice swap on top of the same delivery.
Is this the same as dubbing a video into another language?
No — re-voicing swaps the vocal identity in the same language, keeping the same words and the same timing. Dubbing translates the content into a different language entirely, which requires new timing to match different word lengths and is a separate workflow built for that job.
Keep the take. Change the voice. The performance was never the problem — re-voice it instead of asking a fresh generation to rediscover timing you already had.