Guides

    Your cloned voice sounds nothing like you

    Clone quality is set by the reference recording, not the model. Source-audio requirements and the three defects that guarantee a bad clone.

    Versely Team9 min read

    You can retry the clone twenty times on the same file and you will get the same stranger. Instant cloning does not "learn you" across attempts. It embeds the clip you uploaded. If that clip is noisy, has a second person in it, or has already been through a podcast chain, the embedding is a portrait of the chain. The model is doing its job. The recording is the wrong job.

    Versely’s clone_voice_from_audio is an instant Cartesia clone: a short sample in, a reusable voice_id out, usable with Cartesia generate_speech and change_voice. Cartesia’s own instant-clone guide is the source of truth for that clip: up to ten seconds, speech from start to finish, no extra silence, recorded clearly, in the language you want the clone to speak. A professional fine-tune is a different path (thirty minutes minimum of studio-quality single-speaker audio, two hours or more if you have it). Neither path will rescue a bad reference. Fix the reference, then clone once.

    The recording is the model

    An instant clone is a voice embedding plus the base TTS model. The embedding has to compress whoever is in that clip: pitch, timbre, pace, room, noise floor, and whatever processing is baked into the file. There is no separate "ignore the HVAC" switch. Cartesia is explicit about the consequences:

    • The clip should match the energy you want later. A monotone read of a colourless transcript produces a monotone clone. A clip with the pace and volume of the finished narration produces that pace and volume.
    • Pauses are cloned. Long gaps between sentences become part of the voice. Trim them.
    • Language in the clip is the language of the clone. If you need Spanish, speak Spanish in the sample, or use localisation after. Do not assume an English clip becomes a native Spanish speaker on its own.
    • For a professional clone, speed and volume are fixed from the dataset. You do not get those controls back at request time. Record the pacing you want to live with.

    That is why "I will just pick a better model" is the wrong next move. Best voice cloning models is a useful ranking when you are choosing a path. It will not turn a laptop-fan recording into you. Voice fine-tunes versus instant clones is the comparison of which Cartesia artefact you get. This post is about the tape you feed either of them.

    If you do not actually need a clone, stop here and cast a named voice from the catalog. Cloning is for identity you cannot get off the roster.

    What a usable reference actually is

    Requirements differ by path. Do not mix them.

    Instant clone (Cartesia, and Versely’s clone_voice_from_audio)

    Cartesia asks for a clip of up to 10 seconds. That is a ceiling, not a suggestion to pad. Fill it with speech.

    • One speaker, close, clear, in a quiet room.
    • A script that matches the job: if the clone will narrate product videos, read two or three narration sentences, not a whispered aside and not a shout.
    • No music, no bed, no intro sting.
    • Trim so the file is speech from the first sample to the last. No two-second pre-roll, no tail of chair noise.
    • Export a WAV if you can. A heavily compressed voice-memo is already a processed portrait.
    • Speak in the target language.

    Professional clone / Cartesia fine-tune

    Cartesia’s PVC docs: 30 minutes minimum, 2 hours or more for the better result, studio-quality, single speaker. The training run itself then takes up to about 3 hours. Record at the volume and pacing you want locked in. Instant cloning is clone_voice_from_audio. The agent’s fine-tune tools (list_finetunes, get_finetune_status, list_finetune_voices) list and track a Cartesia fine-tune; they do not start one.

    Either path

    • Mic 15–20 cm from the mouth, pop filter if you have one. A phone in a carpeted room, 15 cm from the mouth, with no speakerphone and no "voice isolation" toggle, is acceptable for an instant clone. A USB condenser in a treated corner is better.
    • Raw capture. No noise suppression, no auto-EQ, no limiter crushing the peaks, no reverb plugin "to make it sound like a studio".
    • One take of clean speech beats a compilation of leftover podcast clips. Leftover podcast clips are how you clone the show’s compressor.

    Then run clone your voice for videos or ask the agent to clone your voice from a recording. Name it something you can find later (Weekly Recap Host, not clone 3). list_cloned_voices is capped at 30 and is easy to clutter with retries of the same bad file.

    Three defects that guarantee a bad clone

    These three survive every retry. If any of them is in the file, do not submit it.

    1. Noise, room, and fan. HVAC, traffic, a laptop fan, a dishwasher, a second room with a TV. The embedding treats that floor as part of the voice. Generated speech then carries a ghost of the room, or a grain that moves with the words. You cannot prompt it out. You cannot EQ it out of the clone. Re-record in a quieter place. A wardrobe full of clothes is a better booth than a treated-looking kitchen with the fridge running.

    2. A second voice, a music bed, or bleed. Instant cloning has no speaker diarisation you can trust for this. Two people on the clip, a YouTube tutorial playing in the background, a jingle at the head of the file, your own voice layered on a track: the embedding is a blend. Generated lines will drift toward the louder source, or smear halfway between them. One speaker. Silence around them. If you can hear another show in the headphones, the clip is already ruined.

    3. Processing. This is the one operators miss, because the file "sounds better" than the raw take. Noise suppression (hard gates, RNNoise-style plugins, Zoom’s "suppress background noise") eats consonants and leaves a watery midrange; the clone inherits the water. Heavy compression and limiting flatten dynamics the TTS model will then reproduce as a squashed, always-on presenter. Reverb, whether a plugin or a live kitchen, becomes a room the clone cannot leave. Clipping is permanent distortion on every sibilant. Speakerphone and Bluetooth headsets add their own bandwidth hole. If the file has already been through a podcast template, it is not a reference. It is a product. Clone the product and you get a copy of the product, not of you.

    A fourth issue is not a defect in the file so much as a mismatch: a 10-second instant clip of you being sarcastic will not give you a calm explainer, and a PVC trained on whispered ASMR will not give you a trade-show booth voice. Cartesia tells you to record the energy you want. Believe that.

    Do not retry the same sample

    Credits on Versely are spent per clone submission. Re-uploading the same WAV is a spend with no new information. The loop that actually converges:

    1. Listen to the source on headphones, loud enough to hear the noise floor. If you hear a fridge, stop.
    2. Re-record. Change rooms, not models. Turn off suppression. Get closer to the mic. Trim the silence.
    3. Clone once.
    4. Test with three lines the series will actually use: a calm statement, a question, a line with a product name. If all three sound like you, keep the voice_id. If they sound like you-in-a-bathroom, the room is still in the clip.
    5. Only then consider a fine-tune, and only if you can produce thirty clean minutes. A fine-tune on noisy podcast exports is a more expensive portrait of the same defects.

    Consent is not optional. Clone voices you own or have written permission to use. Celebrity clips, a coworker who did not sign off, a customer call: those are not reference audio.

    Once the voice_id is good, generate with it through AI voice cloning and the Cartesia speech path. Identity is now the sample's problem, solved. Pronunciation of brand names, number expansion, and long-paragraph flattening are script problems, and they still apply to a perfect clone. The voice cloning glossary is the mechanism; this recording is the input it actually needs.

    FAQ

    How long does the sample need to be?

    For Versely’s instant clone, follow Cartesia’s instant-clone guide: a clip of up to 10 seconds, filled with speech, trimmed of silence. Longer is not better on that path; padding with pauses teaches the clone to pause. A professional Cartesia fine-tune is a different artefact: 30 minutes minimum, 2 hours or more preferred, studio-quality, one speaker.

    The clone sounds like me on the test sentence and like someone else on the script. Why?

    The test sentence was close to the reference read (same pace, same pitch). The script is a different performance: faster, louder, a question, a product name. Record a reference that already contains that range, still inside the instant-clip length, or accept that an instant clone is a snapshot of one delivery. A PVC is the path when you have the minutes and need the range.

    Can I clean the audio with iZotope or a "denoise" button and then clone?

    You can, and you will often clone the denoise. Light, inaudible cleanup of a single click is fine. Aggressive suppression, voice isolation, and "podcast ready" presets are defect 3. Re-record. It is faster than trying to invent a dry signal from a processed one.

    Does picking a different cloning model fix a bad clip?

    No. Instant clones and fine-tunes both consume the reference. A cleaner model on a noisy file is a cleaner picture of the noise. Change the recording. If you do not need identity, skip cloning and use a named catalog voice.