Create · Versely Agent

    Design a custom AI voice from a description

    Describe the voice. Skip the preset list.

    What you say to the agent

    No special syntax — just describe it like you would to a person.

    Design a gravelly older narrator with a slow, confident cadence and read this intro
    Give me a bright, upbeat young voice with a slight British accent for this line

    What it does, step by step

    1. 1

      Describe the voice character in words instead of picking from a fixed voice_id list.

    2. 2

      Give it the text you want spoken in that voice.

    3. 3

      It designs the voice (Qwen 3 TTS Voice Design) and speaks your text with it in one call.

    What it needs from you

    • The text/script you want spoken
    • A text prompt describing what you want

    What comes back

    One audio clip in the described voice.

    What it costs

    Priced per generation, shown before you confirm.

    Good to know

    • This doesn't create a reusable voice_id for other tools — re-describe the same voice to get a similar sound on new text. For a reusable voice, use voice cloning from an audio sample instead.

    Under the hood

    This is what the agent actually calls when you ask for it — real tools from its live surface, not marketing copy.

    create_custom_voice

    Design a custom voice from a text style prompt (Qwen 3 TTS Voice Design) and speak the given text with it in one call. Unlike generate_speech's fixed voice_id list, the voice itself is described in words (e.g. 'a gravelly older narrator with a slow, confident cadence'). Returns one audio clip synchronously — this does not create a reusable voice_id for other tools; re-run with the same prompt for a similar-sounding voice on new text.

    The full tool behind it

    See it done in a real workflow

    Or start from a one-tap template

    Frequently asked questions

    What do I actually say to the agent to design a custom AI voice from a description?+

    Just describe it in plain English — for example: "Design a gravelly older narrator with a slow, confident cadence and read this intro" The agent handles picking the right tool and model from there.

    What does the agent need from me first?+

    At minimum: The text/script you want spoken; A text prompt describing what you want. Anything else it needs, it asks for before running.

    What do I get back?+

    One audio clip in the described voice.

    Does this cost credits?+

    Priced per generation, shown before you confirm.

    Anything I should know before asking for this?+

    This doesn't create a reusable voice_id for other tools — re-describe the same voice to get a similar sound on new text. For a reusable voice, use voice cloning from an audio sample instead.

    You can also just ask for

    Ask your Versely agent to design a custom AI voice from a description

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.