Guides

    Cartesia Voice Clone: a voice file, not a talking generate (8cr)

    Cartesia Voice Clone writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Versely Team3 min read

    Cartesia Voice Clone writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Cartesia Voice Clone is Cartesia's voice-clone row: 8 credits, audio out, no picture, no duration list. The catalog job is "Clone any voice using Cartesia AI. Upload a short audio sample to instantly create a personalized voice for text-to-speech generation." Features listed: Voice Clone, Custom Voice, Audio Upload, Multilingual, High Quality. That file is a track. It is not a talking head.

    Upload a sample, get a voice

    The input is a short audio sample. The output is a voice you can drive with text. The text-to-speech tool is the door. Voice cloning is the mode. Nothing in that loop opens a mouth on screen.

    If the cut is a face that has to speak, this row is the voice pass only. The mouth pass is lipsync or a talking generate. Laying a cloned track under a still face is how the audience reads a dub they were not supposed to notice. Add a voiceover when we do not see the jaw. AI lipsync when we do.

    8 credits is the clone, not the clip

    Eight credits is the catalog price for this generate. There is no resolution because there is no frame. There is no 4s / 8s ladder because this is not video. Plan the script, clone the voice, then pick the picture row. Do not buy an 8-credit clone hoping the waveform will invent lips.

    Multilingual is a listed feature. That is useful for a localized VO. It is not a translated talking-head pipeline by itself. Translation plus a closed mouth is still a closed mouth. The dubbing, lipsync, and cloning guide is the stack. This page is only Cartesia's clone row.

    Where the file goes next

    Three honest landings:

    • Off-screen narrator: clone, generate the line, lay it under b-roll.
    • On-screen speaker you already shot: clone or replace the voice, then run a lipsync row on the plate.
    • Character who does not exist yet: do not start here. Start with a picture. Come back for the voice if the picture is a still you will later drive.

    Clone your voice for videos is the editing-task version of the same split. Cartesia Voice Clone is 8 credits and a sample. It will not fix a face that is not speaking.

    Do not lay this on a closed mouth

    The test is a still frame from the cut. If the mouth is in shot, this file is incomplete delivery. If the mouth is not in shot, this file is the delivery. Stay honest about which cut you have. Eight credits of cloned audio on a silent jaw is not a talking generate. It is a mismatch.

    FAQ

    Does Cartesia Voice Clone make video?

    No. Content type is audio. Category is voice-clone. You upload a sample and get a voice for text-to-speech. There is no picture.

    Do I need an image?

    No. This row does not require an image. An image would not help. The mouth lives on a different row.

    Is this the same as lipsync?

    No. Lipsync drives a mouth from a track. This row writes the track. If we see the mouth, run lipsync after, not instead of thinking.

    What are the 8 credits for?

    The catalog price for this clone generate. Not a 4K talking clip, not a duration ladder, not a scene.