Speech, voice and audio

    Voice design

    Voice design creates a new synthetic voice from a written description — age, accent, texture, energy — instead of cloning one from a recording.

    There is no source speaker. You describe the voice you want in the same way you would brief a casting director, and the model synthesises something matching it. That makes the output nobody's voice in particular, which sidesteps the consent question that cloning necessarily raises.

    It is the right tool for characters and brand voices that should not be traceable to a person: a narrator for a faceless channel, a mascot, a set of distinct voices for an animated cast. It is the wrong tool when the point is that a specific real person is speaking.

    Descriptions behave like prompts, with the same quirks. Concrete, physical attributes — pitch, pace, warmth, breathiness, accent — steer reliably; abstract adjectives like "trustworthy" mostly do not. And a designed voice is only as reusable as your ability to reproduce it, so save the description alongside the result.

    In practice

    • Describe physical qualities, not personality traits.
    • Save the exact description — it is the only record of how the voice was made.
    • Generate a few candidates and audition; description-to-voice is not one-to-one.

    The mistake to avoid

    Designing a voice per project. Distinct channels need distinct voices, but one channel needs one voice used consistently for months.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.