Speech-to-speech keeps the performance; multilingual TTS invents one
Speech-to-speech is a recording in, a new timbre out — timing and emphasis survive. ElevenLabs Voice Change is audio-to-audio at 5 credits. ElevenLabs Multilingual is text-to-audio at 6 credits. Do not feed a script to the changer.
/glossary/speech-to-speech is a recording in, a new timbre out — timing and emphasis survive. ElevenLabs Voice Change is audio-to-audio at 5 credits. ElevenLabs Multilingual is text-to-audio at 6 credits. Do not feed a script to the changer.
The glossary already owns the definition. This page owns the fork in the catalog: which ElevenLabs row matches the file you actually have. If you have a take, convert it. If you have words, synthesise them. Treating those as the same button is how a good read gets thrown away and how a script gets stuffed into a converter that has nothing to convert.
The glossary term is a recording in
Speech-to-speech takes a recording of one person talking and re-renders it in a different voice, keeping the original performance intact. That sentence is the whole product boundary. Timing, emphasis, pauses, the laugh in the middle of a line survive because you acted them. The converter swaps timbre. It does not invent a delivery.
Text-to-speech is the other boundary: written text in, a synthetic voice out. The engine has to guess where the stress goes, how long to hold a beat, whether the line is a joke. Plain text does not contain a performance. Two takes of the same script are two invented readings. That is a feature when you are casting. It is a loss when you already directed the take.
Do not rewrite the glossary as a prompt knob. There is no slider that turns a paragraph into a performance. The input type is the knob: recording versus script.
The ceiling is the source. Conversion carries over what it hears. A mumbled line stays mumbled. A flat read becomes a flat read in a different timbre. If the complaint was the performance, conversion is the wrong tool.
ElevenLabs Voice Change is audio-to-audio at 5 credits
ElevenLabs Voice Change is the catalog row for that conversion. Content type audio. Category audio-to-audio. 5 credits. Catalog audio true. No resolution. No duration enum. You start with a recording. You get a recording. You do not get a face.
Use it when the words and the timing are keepers and the voice is not. Placeholder narration swapped for a roster voice. A temp read swapped for a cloned timbre you already have consent to use. Off-camera dialogue that should not be re-typed.
It will not invent a presenter or a plate. It will not lip-sync a mouth. If we see the face, a converted file laid on a closed mouth is still a closed mouth. Send the new track into a lipsync job, or convert audio that was always meant to live off camera.
Change the voice in a video is the picture-side cousin: replace the bed on a clip. Voice Change is the audio-to-audio engine. Trim before you convert. Keep the original. Conversion is a generate, not an edit.
ElevenLabs Multilingual is text-to-audio at 6 credits
ElevenLabs Multilingual is multilingual text-to-speech powered by ElevenLabs Multilingual v2. Content type audio. Category text-to-audio. 6 credits. No still. No duration enum. You start with a script. You get a performance the model invented.
That is the correct row when there is no take: a new VO, a language pass from copy, a line nobody has spoken yet. Punctuation is prosody. Write numbers the way they should be read. Audition more than one take; the output is stochastic.
It is the wrong row when the pause before the punchline already exists as a file. Re-typing that line into Multilingual throws the pause away and asks the engine to rediscover it. Sometimes it will. Often it will not. The 6-credit sticker is not cheaper conversion. It is a different input.
AI text to speech is the tool shape for that family: engines, not just a voice dropdown. Multilingual is one catalog row on that shape. Voice Change is not on that shape. Do not open TTS because the converter felt obscure.
Do not feed a script to the changer
The changer has nothing to convert if you paste words. A script is not a recording. Feeding copy into an audio-to-audio row is a mislabel, the same as uploading silence and hoping for a read.
The intake is five checks:
- Is the delivery actually right? If not, recut or re-record. Then convert.
- Is the vocal isolated? A music bed baked into the take comes along for the ride.
- Is there one speaker? Two voices overlapping are not one target.
- Are the lips on camera? Conversion does not re-time a mouth.
- Do you have the right to put those words in that voice?
Four yeses and a clean consent answer means Voice Change. A script with no file means Multilingual. A close-up talking head means lipsync after, or a reshoot.
Do not use Voice Change to "fix" a bad read. It changes who is speaking, not how well the line was delivered. Do not use Multilingual to keep a performance you never recorded. Same vendor family. Different content type. Pick the row that matches the input you actually have.
FAQ
Can I paste a script into ElevenLabs Voice Change?
No. Voice Change is audio-to-audio at 5 credits. It needs a recording. A script belongs on ElevenLabs Multilingual at 6 credits, text-to-audio.
Does conversion rewrite the words?
It is not supposed to. Speech-to-speech preserves speech content and swaps timbre. If you need different words, record them or run text-to-speech. Do not hope the changer will ad-lib.
How is this different from the glossary page?
The glossary defines the term. This page is the catalog fork: Voice Change versus Multilingual, 5 credits versus 6, recording versus script. Do not treat the term as a prompt knob.
If the mouth is on camera, which row do I pick?
Neither row writes a talking-head video. Convert or synthesise the audio, then lipsync the plate, or shoot a real person. Laying a new timbre on a closed mouth is still a closed mouth.