Audio, Voice & Dubbing · Versely AI

    Clone Your Voice for Videos

    Record once. Reuse the voice forever.

    clone_voice_from_audio takes a sample recording and returns a voice_id — a reusable Cartesia voice you can call on for every future generate_speech or change_voice job, instead of picking from the stock catalog or re-recording each time.

    This is the foundation for scaling narration in your own voice across many videos without re-recording, and for keeping voice identity consistent across a series.

    Powered by

    clone_voice_from_audiogenerate_speechchange_voicelist_cloned_voices

    Clone a reusable Cartesia voice from an audio sample. Returns a voice_id usable with generate_speech (Cartesia models) and change_voice. Use when the user wants to create a voice from a recording rather than pick from the existing catalog.

    What to tell the agent

    Versely's agent maps this plain-English request directly onto clone_voice_from_audio. You don't need to know the parameter names — just describe what you want.

    Clone my voice from this audio sample and name it 'Main Narrator', then use it to read this script.

    How it works

    1. 1. Provide a clean sample

      Upload or point to an audio_url with clear, representative speech — the cleaner the sample, the better the clone.

    2. 2. Name and describe it

      name is required; description and language are optional but help you tell voices apart later.

    3. 3. Get your voice_id back

      The returned voice_id is what you pass to generate_speech (Cartesia models) for narration, or change_voice to re-voice existing recordings.

    4. 4. Reuse across projects

      Browse your saved voices anytime with list_cloned_voices instead of re-cloning for every new video.

    What it costs

    Cloning is billed once per sample submitted; reusing the resulting voice_id in generate_speech is billed at that tool's normal per-generation rate. Confirm exact pricing with estimate_cost.

    The formula behind that number: What does an AI voiceover cost for my script? meters the characters you type, not the seconds you get back.

    Limits & things to know

    • Quality depends heavily on the sample's audio cleanliness — background noise or multiple speakers in the sample will degrade the clone.
    • The cloned voice is only usable with Cartesia-backed calls (generate_speech with a Cartesia model, and change_voice) — it isn't a universal voice_id across every TTS model in the app.
    • Only clone voices you own or have explicit consent to use.

    Who uses this

    • Narrating a whole video series in your own voice without re-recording
    • Personal-brand YouTube channels
    • Consistent voice identity across a UGC ad campaign
    • Preserving a voice for future dubbing or narration work

    Frequently asked questions

    How much audio do I need to clone a voice?+

    A clean, representative sample is enough — clone_voice_from_audio doesn't require a long recording, but clarity matters more than length.

    Where can I use the cloned voice afterward?+

    The returned voice_id works with generate_speech (on Cartesia models) for new narration, and with change_voice to re-voice an existing recording.

    Can I see all my cloned voices later?+

    Yes — list_cloned_voices returns your saved voices so you can reuse one without cloning again.

    Can I clone someone else's voice?+

    Only clone voices you own or have explicit consent to use — third-party or celebrity voices require verification and disallowed use cases are blocked.

    Related Versely tools

    Related editing jobs

    Clone Your Voice for Videos inside Versely

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.