Guides

    Designing a custom AI voice returns a clip, not a voice you can reuse

    create_custom_voice is Qwen voice design plus one read. Cloning a reusable voice_id is a different job, and designing every line wastes credits.

    Versely Team4 min read

    Describe the voice. Skip the preset list. Get one audio clip. That is the job. Design a custom AI voice is one named agent capability. create_custom_voice designs the voice from a text prompt (Qwen 3 TTS Voice Design) and speaks your text in the same call. It does not mint a reusable voice_id. Re-running the same description on new copy gives you something similar, not something identical. Treating design as a brand voice — and paying to re-describe it on every line — wastes credits on a character that will not hold.

    Voice design is a casting pass for a single spot. If you need the same throat next week, that is a different named job.

    One description, one read, one file

    You pass two required fields: text (the words to speak) and prompt (the voice, in words). Example shape: "a gravelly older narrator with a slow, confident cadence." Optional language and sampling knobs sit beside that. The tool returns one audio clip synchronously. There is no roster row waiting for you afterward.

    That is the compliment. You are not scrolling a voice_id list hoping "warm British woman 4" is close. You write the register. You get a take.

    It is also the limit. The clip is the artifact. The description is not a saved identity. Ask the agent to "use that voice again" on a second paragraph and you are asking it to roll another design that sounds like the first. For a one-off character, similar is enough. For a series host, similar is drift you will hear by episode four.

    Cost is priced per generation and shown before you confirm. The unit is the clip, not a voice you now own.

    What this job is not

    A clone from a recording. Clone my voice from a recording is clone_voice_from_audio. It needs a sample, a name, and it returns a voice_id you can call from speech and voice-change tools later. Design has no sample. Clone has no "gravelly narrator" paragraph. Pick one.

    A preset read. Write and generate a voiceover is how you pick a catalog voice and speak a script. If a roster voice is already right, do not design a cousin of it. Design is for the register the list does not have.

    A re-voice of an existing take. Change the voice in an audio clip keeps delivery and timing and swaps the throat. Design speaks your typed text in a newly described voice. It will not preserve a performance you already recorded.

    The voice cloning tool is the family door. This capability is only the describe-and-speak row.

    Spend design credits on a one-off character

    The waste is using create_custom_voice as a cheap clone. You like the first read. You generate the next three lines "in the same voice." You now have four related strangers. A weekly show built that way is four hosts.

    If the voice has to survive a season, clone it from a take you actually like — or pick a catalog id and stop. If the voice is a one-shot villain, a temp narrator, a character who appears once, describe it, speak the line, ship the clip. Do not "save" a design that cannot be saved.

    Do not put the spoken words in the style prompt, and do not put the timbre in the text field. Two fields, two jobs inside one call. Mixing them is how you get a prompt that tries to say the line.

    FAQ

    Does create_custom_voice give me a voice_id I can reuse?

    No. The tool returns one audio clip. Re-describe the same voice for new text if similar is enough. For a reusable identity, clone a voice from a recording.

    Can I design a voice once and then generate_speech with it?

    Not as a saved id. generate_speech wants a voice_id from the roster or from a clone. Design does not write into that list. Asking the agent to "use the designed voice in TTS" is how you accidentally roll another design — or pick a preset and lose the character.

    When should I pick a catalog voice instead?

    When the register already exists on the list and you need consistency more than novelty. Design is for the gap: gravel, age, cadence, accent the presets do not cover, on a line you will speak once.

    Why not clone every designed clip I like?

    You can clone a take after you hear it, if that take is the identity you want. That is a second named job with a sample. Do not skip the sample and hope the description is the clone. Description is not audio.