ElevenLabs Voice Change: a voice file, not a talking generate (5cr)
ElevenLabs Voice Change writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
ElevenLabs Voice Change writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
The catalog: voice conversion that transforms voice characteristics while preserving speech content. Content type audio. Category audio-to-audio. 5 credits. Catalog audio true. No resolution. No duration enum. No image required. Features: voice_conversion, voice_cloning, real_time_processing. Billing is per second (matrix floor 25, ceiling 600).
You start with a recording. You get a recording. You do not get a face. The words stay; the timbre changes.
Conversion keeps the words
ElevenLabs Voice Change is speech-to-speech, not text-to-speech. If you only have a script, you are on ElevenLabs Multilingual (6 credits, Multilingual v2, a file from text). If you have a take you like — pacing, breaths, the joke's timing — and the voice is wrong, this is the row.
The change the voice in a video job is the picture-side cousin: strip or replace the bed on a clip. This card is the audio-to-audio engine. AI voice cloning is how you get a target timbre to convert into. AI text to speech is the wrong door if the performance already exists as a file.
Do not re-type the line into TTS because conversion felt obscure. Retyping throws away the take. Conversion is how you keep it.
Five credits is an audio pass
The catalog figure is 5 credits, per second of source. A long interview is a long bill (up to 600 on the matrix). A 5-second hook is the cheap version of this product. Either way you are not buying 720p. You are buying a new voice on old words.
If we see the mouth, a converted file laid on a closed-mouth generate is still a closed mouth. Send the new track into AI lipsync, or convert the audio under a plate that was always meant to speak. If the picture already spoke in a different voice, replacing the bed without a lipsync pass is how the face chews the wrong syllable.
Voice Change is not a talking generate. It will not invent a presenter, a room, or a 4K plate. It will not fix a bad read — it will re-voice it. If the pacing is wrong, recut or re-record, then convert.
Mouths belong on lipsync
Use this row when:
- The words and timing are keepers.
- The timbre is not.
- The picture is mute, off-camera, or about to be lipsynced to this new file.
Do not use it when:
- You have no source audio. That is TTS.
- You wanted a new scene. That is video.
- You wanted the on-camera person to look like they said the new timbre without a lipsync step. That is a missing job, not a missing prompt.
Audio-to-audio is a thin category in the catalog. This card is that category. Treat it like a mix move, not a generate.
FAQ
Can Voice Change output a talking-head video?
No. It writes audio. Preserve the speech content, change the voice characteristics, then — if we see the mouth — lipsync the file to a plate.
How is this different from ElevenLabs Multilingual at 6 credits?
Multilingual starts from text. Voice Change starts from a recording. Same vendor family, different content type: text-to-audio versus audio-to-audio. Pick the one that matches the input you actually have.
Does conversion rewrite the script?
It is not supposed to. The catalog says it preserves speech content. If you need different words, record or TTS the new line. Do not hope a voice-change pass will ad-lib.
Why is the bill per second if the title says 5cr?
5 credits is the catalog rate on the card; the matrix bills that rate per second of audio, with a 25–600 credit window. Length of the source is the length of the bill. Trim before you convert.