Guides

    Cast a named voice before you clone one

    A 210-voice named library covers most briefs without consent paperwork. Here is an audition process and the line where cloning becomes necessary.

    Versely Team9 min read

    Most voice projects start at the wrong end. The brief says "we need a warm, trustworthy narrator for the product explainer series," and within an hour someone is recording a sample, chasing a release form, and building a cloned voice — for a job where a named voice off the roster would have been approved in twenty minutes and would never need a consent conversation at all.

    Cloning is a real capability and it solves a real problem. It is just not the first move. The named library covers 210 voices across 31 languages, and it takes an audition, not a project, to find out whether one of them fits. Doing the audition first is the cheapest way to discover that you do not need the harder path.

    What the named library actually holds

    The roster is deep in some languages and thin in others, and knowing which side your brief sits on decides how long casting takes.

    Language Named voices
    English 80
    Spanish 20
    Korean 15
    Japanese 14
    Portuguese 12
    Italian 7
    French 6
    Indonesian 6
    Russian 4
    German 3
    Hindi 3
    Arabic 2

    Every language below Arabic on that list carries exactly two, which is the shape of the long tail. Thirty-one of the 38 languages carry voice entries and seven carry none at all. Of those 31, fifteen list only generic Male Voice and Female Voice placeholders, which leaves 16 languages with described archetypes you can cast against. English alone accounts for 80 of the 210, which is why English casting is a shortlisting problem and Arabic casting is a yes-or-no problem.

    The names describe casting intent rather than voice IDs, which makes shortlisting fast if you read them as a director would. The English roster includes entries like Expressive Narrator, Trustworthy Man, Gentle Teacher, Mature Boss, Magnetic Voiced Man, Reserved Young Man and Husky MetalHead. Spanish carries Captivating Storyteller, Wise Scholar, Confident Woman and Comedian. Korean carries Intellectual Senior, Calm Gentleman, Brave Female Warrior. These are archetypes, and a brief that says "warm, trustworthy narrator" maps onto three or four of them immediately.

    Sitting alongside the per-language rosters are MiniMax's style buckets: 17 Universal entries (Wise Woman, Deep Voice Man, Calm Woman, Casual Guy, Patient Man and similar) and 4 Character ones — Angry Pirate, Kind Troll, Movie Trailer, Peace & Ease. These are style buckets rather than languages, and they are worth knowing about because a brief for a trailer read or a comedy character often lands on one of them directly.

    The audition, as a process

    Six steps, roughly twenty minutes, and every generation in it draws credits like any other — so keep the script short and the shortlist tight.

    1. Write the audition script from the real copy. Twenty to thirty seconds, pulled from the actual project, not a generic pangram. Include the hardest line: the product name, the number, the sentence with the awkward clause. A voice that handles a clean sentence beautifully and mangles your brand name is not a fit, and a generic sample will never surface that.
    2. Shortlist three to five names on the axis that matters. Pick the axis before you browse — authority, warmth, energy, age impression — and filter on it. Auditioning fifteen voices produces fatigue, not a decision.
    3. Generate all of them on the same text with the same settings. Same emotion or style instruction, same language, same script. A comparison where one voice got a better delivery direction is not a comparison of voices.
    4. Listen on the target device. Short-form audio is phone-speaker audio. A voice that reads as rich on monitors can read as boomy and indistinct through a phone, and the difference between two candidates often inverts between the two.
    5. Stress-test the winner on the worst line in the script. Numbers, acronyms, foreign proper nouns, and any sentence with a clause that changes meaning depending on where the stress lands. This is where a shortlist of three collapses to one.
    6. Record the voice ID somewhere the next person will find it. The casting decision is worth more than the audio it produced. Losing it means re-auditioning in three weeks and shipping a series where episode four sounds like a different person.

    Writing and generating a voiceover is the direct route for steps 3 through 5, and the voice-over hub collects the per-language surface — English has its own dedicated page, as do Spanish, French and Korean.

    Where the library actually runs out

    Three failure conditions, and they are worth checking before the audition rather than during it.

    The language has no castable roster. Seven of the 38 carry no voice entries at all, and 15 more list only generic Male Voice and Female Voice placeholders. You can still generate speech; you just cannot cast against a described archetype, and comparing candidates gets harder because the names stop carrying information.

    The roster is deep but the specific register is missing. Eighty English voices is a lot and it is not infinite. A specific regional accent, a particular age impression, or an unusual texture may simply not be represented, and no amount of emotion direction conjures it — delivery settings shift how a voice performs, not what it sounds like.

    The voice needs to be a particular real person. This one is not a library limitation at all. It is a different requirement, and the library was never going to satisfy it.

    The ladder: named, designed, cloned

    Three rungs, and the cost of each is not primarily credits — it is the overhead you take on.

    Named voice. Pick from the roster, note the ID, done. No consent paperwork, no source recording, no reproducibility risk. The voice is stable, it is there next quarter, and anyone on the team can use it by ID.

    Designed voice. Describe the voice in words instead of picking one, through Qwen's voice-design path. Useful when the roster has nothing in the register you want and the voice is for a one-off character rather than a series. The constraint worth taking seriously: this produces a clip, not a reusable voice identity. Re-running the same description on new text gives you something similar, not something identical. It is a casting solution for a single spot, not a brand voice. Designing a custom AI voice covers how it works and what goes in which field.

    Cloned voice. Build a reusable voice from an audio sample, get back an ID that works with the speech tool and with voice conversion. This is the rung with real overhead — a clean source recording, consent from the person whose voice it is, and a policy for how the identity is stored and who may use it. It is also the only rung that solves the problems below.

    When cloning is genuinely required

    Four cases. If your project is not one of them, the ladder stops at rung one or two.

    • The voice is a specific real person and that is the point. A founder narrating their own story, a presenter with an established audience relationship, an actor under contract for the campaign. There is no roster substitute for identity.
    • The same identity has to persist across many future scripts. A named voice does this too, but if the identity is a person rather than an archetype, cloning is the only route that keeps them consistent across content that does not exist yet.
    • A dub has to preserve the original speaker. Voice-preserving dubbing is the case where the source speaker's identity carries into other languages. That is a cloning operation by construction.
    • You need to re-voice existing takes into that identity. A cloned ID works with voice conversion as well as with text-to-speech, which lets an existing recorded performance be re-voiced rather than re-performed. A named roster voice does not give you that chain in the same way.

    If you are on this path, two things are worth reading first: the difference between an instant clone and a fine-tune, covered in voice fine-tunes versus instant voice clones, and what cloning actually is as a category. The voice cloning tool page is the entry point once the consent side is settled.

    One operational note for teams that clone often: the cloned-voice listing returns your voices newest first and caps at 30. That is plenty for a working set and not plenty for an agency that clones a voice per client and never prunes. Name them properly and retire the ones that are done.

    FAQ

    How long should an audition script be?

    Twenty to thirty seconds of real project copy. Long enough to include your hardest line — a brand name, a number, an awkward clause — and short enough that generating five candidates is a quick pass rather than a session. Generic sample text is worse than useless because it hides exactly the failures you are auditioning to find.

    Can I get a named voice to sound older, angrier, or more energetic?

    Within limits. Delivery settings — an emotion parameter or inline tags, depending on the engine — change how a voice performs a line. They do not change the underlying timbre or age impression. If the brief needs a fundamentally different voice, that is a casting change, not a direction change.

    Is a designed voice a substitute for a cloned one?

    Only for one-offs. Voice design produces a clip from a written description rather than a reusable identity, so running the same description again gives you something in the same territory rather than the same voice. For a character in a single spot that is fine. For a series where episode nine has to match episode one, it is not.

    What if my target language only has two or three named voices?

    Audition all of them, since the shortlist is already the whole roster, and check the voice-over hub for what the language supports before promising a specific read. If none fit and the language falls outside the cloning map, the honest answer is that the register you want may not be available yet — plan the script around the voices that exist rather than around one that does not.