Speech, voice and audio

    Voice cloning

    Also called Voice replication, Custom voice.

    Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.

    Two levels exist in practice. A quick clone works from a short sample — often under a minute — and captures timbre convincingly while being looser on the speaker's habits of rhythm and emphasis. A high-fidelity clone wants substantially more audio and gets closer to a specific person's delivery, not just their tone.

    Sample quality dominates everything else. One speaker, one room, no music, no overlap, consistent distance from the microphone, and enough range that the model hears more than one emotional register. Ten clean minutes beat an hour of podcast audio with a co-host in it, because the model treats whatever is consistently present as part of the voice.

    The obvious caveat is not a technicality: a clone of someone else's voice needs their permission, and both the platform's terms and the law in most jurisdictions treat it that way. Recorded consent from the speaker is the norm for commercial use.

    In practice

    • One speaker, clean room, no background music — everything else compromises the clone.
    • Include varied delivery in the sample: the model can only reproduce registers it heard.
    • A clone is a reusable asset — build it once carefully rather than repeatedly in a hurry.

    Voice cloning models

    Catalog entries that build a reusable voice from a sample you provide. 5 of the 296 models in the Versely catalog qualify.

    ModelProviderType
    Cartesia Sonic 3.5CartesiaAudio
    Qwen 3 TTS Voice DesignQwenAudio
    Cartesia Voice CloneCartesiaAudio

    The mistake to avoid

    Cloning from a video's existing soundtrack. Ambience, music and room reverb are learned as part of the voice and turn up in every line it later speaks.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.