AI Models

    Designing a Custom Brand Voice With Qwen 3 TTS

    How Qwen 3 voice design lets brands build a custom TTS voice from a text description: attribute prompting, iteration workflow, and design vs cloning.

    Versely Team7 min read

    "Warm, unhurried, mid-30s, slight rasp, sounds like she is explaining something she genuinely finds interesting." That sentence is not a casting brief I sent to a voice agency. It is, almost verbatim, the prompt that produced the narration voice a brand I work with now uses on every video. No auditions, no session fees, no human whose availability or exit could ever complicate things: the voice was designed, not cast.

    Voice design is the third path in synthetic speech, after stock voices and cloning, and Qwen 3's voice design capability inside Versely is the cleanest implementation of it I have used. You describe the voice you want in plain language and the model builds it. This post covers how to write voice descriptions that work, the iteration loop, and when design beats cloning.

    Sound mixing board with rows of faders in a recording studio

    The three paths to a brand voice, quickly

    • Stock voices are excellent and instant, but shared: the voice fronting your brand is fronting hundreds of others.
    • Cloning reproduces a specific real person, which is exactly right when that person carries brand equity, and brings consent contracts and human dependencies with it; the full setup and ethics are in voice cloning for brand narration.
    • Design synthesizes a voice that has never existed, from a description. No person, no consent overhead, no exclusivity question: the brand owns a voice nobody else can be, because nobody else is it.

    Design is the underused option, mostly because until recently the results sounded like slightly randomized stock voices. Qwen 3's descriptive control changed that: the attributes you specify genuinely land in the output.

    Writing a voice description that works

    Treat the description like a character brief, not a settings panel. From a lot of iterations, the attributes that reliably steer Qwen 3, in descending order of impact:

    Attribute Example phrasing Effect strength
    Age and gender "woman in her early 40s" Strong and reliable
    Pace "unhurried, leaves small pauses" Strong
    Texture "slight rasp," "smooth and rounded," "breathy" Strong
    Energy register "calm authority," "bright and eager" Moderate to strong
    Attitude "sounds like she finds this genuinely interesting" Moderate, but the magic when it lands
    Accent flavor "neutral American," "soft British inflection" Moderate
    Metaphor references "late-night radio host," "documentary narrator" Variable, but efficient shorthand

    Two composition rules matter more than any single attribute. First, describe a person, not a sound file: "a mentor who has explained this a hundred times and still enjoys it" outperforms a list of adjectives. Second, include what the voice is not when you keep getting an unwanted quality: "not salesy, not announcer-like" prunes the default TTS enthusiasm that infects everything.

    A worked example from a real project, a DTC sleep brand: "Man in his late 30s, low and quiet without whispering, very unhurried, warm, neutral American accent, sounds like the last calm person awake. Not dramatic, not an announcer." Third iteration was the keeper.

    The iteration loop

    Voice design is a convergence process. My loop, which usually lands inside five rounds:

    1. Write the brief and generate against a fixed test script. Use the same 60-word script every round: one question, one list, one emotional beat, one product name. Changing the script between rounds destroys your ability to compare.
    2. Diagnose in one dimension at a time. Too fast and too formal? Fix pace first, regenerate, then address formality. Multi-attribute edits make it impossible to know what moved the result.
    3. Push one attribute past where you want it, then pull back. Descriptions are interpreted conservatively; "very slow" often produces "slightly slow."
    4. Test the finalists on real scripts, especially your product names and any jargon. A voice that aces the test script can stumble on your actual vocabulary.
    5. Panel-check with three people who did not write the brief. The designer hears intent; fresh ears hear the voice. Ask them to describe it in three words and check the words against the brand book.

    Once locked, the design brief itself becomes a brand asset. Document it next to your logo files: it is the reproducible recipe for your audio identity.

    Keeping a designed voice consistent in production

    A designed voice is only a brand voice if it sounds identical across every video, and consistency in TTS is a discipline:

    • Lock the voice and reuse it; never regenerate from the brief per project. The brief is the recipe, but the saved voice is the product. Regenerating produces siblings, not twins.
    • Standardize script punctuation conventions, because punctuation is your pacing control. A house rule as simple as "dashes for beats, paragraphs for topic shifts" keeps different writers producing the same delivery.
    • Maintain the pronunciation glossary for names and jargon, same as you would with a clone.
    • Guard the register. A calm designed voice reading an ALL-CAPS hype script produces an uncanny mismatch. Either the script conforms to the voice or the content belongs to a different format.

    Consistency techniques for keeping synthetic delivery stable across long content runs are covered in more depth in style-locking AI text-to-speech, and they apply directly here.

    The deployment surface is everything Versely produces: narration over generated footage in the AI video generator, UGC ad voiceovers, slideshow narration, and dubbed variants where the designed voice carries into other languages.

    Design vs. clone: the actual decision

    After running both paths for different brands, the decision compresses to one question: is there a human whose voice already means something to your audience?

    • Yes, and they are committed to the brand: clone. Equity beats novelty.
    • Yes, but the relationship is uncertain: design. A voice entangled with a departing founder or an expiring narrator contract is a liability with a countdown.
    • No existing voice: design, almost always. You skip consent contracts, session logistics, and exit clauses entirely, and you get a voice fit to the brand book instead of fit to whoever was available.

    There is also a hybrid I have started recommending: design the brand's everyday narration voice, and reserve the founder's cloned voice for moments where personal authorship matters. Two voices, two jobs, no confusion about which is speaking.

    FAQ

    What is voice design in Qwen 3 TTS?

    It is generating a new synthetic voice from a natural-language description of its characteristics: age, pace, texture, energy, accent, and attitude. The output is a reusable voice that never belonged to any real person, which the brand can deploy across all its content.

    How is voice design different from voice cloning?

    Cloning reproduces a specific real person's voice from recordings and requires their consent and a contract. Design synthesizes a voice from a description with no source human, so there are no likeness rights, no consent overhead, and no dependency on a person's continued involvement.

    How many iterations does it take to get a good brand voice?

    With a fixed test script and one-attribute-at-a-time edits, usually three to five rounds. The slow path is changing many attributes at once, which makes results impossible to attribute. Budget an afternoon for the design session and a panel check.

    Can a designed voice be used commercially?

    Yes; designed voices carry none of the likeness questions cloning raises, since no real person's voice is reproduced. Standard platform disclosure rules for synthetic media still apply where required. Avoid deliberately designing an imitation of a specific real person, which reintroduces every problem design exists to avoid.

    Does a designed voice work for ads as well as narration?

    Yes, but design it for the job: ad reads want more forward energy than tutorial narration. Many brands design two registers of the same character. How voice choice affects ad performance specifically is covered in text-to-speech for video ads.

    Write the one-sentence casting brief for the voice your brand should have had all along, and run the design loop in AI voice cloning. Keep the brief; it is the recipe. Free credits daily.