Comparisons

    Seven voice engines and how to pick one

    Language reach, cloning, expressiveness and tag handling separate the seven engines in Versely's voice table. Here is a default engine per use case.

    Versely Team9 min read

    Engine choice for speech gets argued as a taste question and decided as a constraint question. Seven engines carry Versely's per-language voice table — MiniMax, ElevenLabs, Inworld, Gemini, Grok, Cartesia and Qwen — and for most real jobs, two of the four axes below eliminate five of them before anyone plays a sample. Knowing which axis binds first saves the afternoon you would otherwise spend auditioning voices in an engine that does not support your target language.

    The four axes, in the order they usually bind: language reach, how you direct emotion, whether cloning is available, and the structural features that only exist on one engine each.

    Axis 1: language reach, which is not the same as voice choice

    The current language table covers 38 languages across the seven engines. They do not cover them evenly.

    Engine Languages covered
    MiniMax 31
    ElevenLabs 30
    Gemini 24
    Grok 16
    Cartesia 15
    Inworld 15
    Qwen 10

    Read that as a gate, not a ranking. If the brief says Slovak, the list is ElevenLabs and nothing else. Bulgarian and Croatian are the same. Hungarian, Norwegian and Persian go the other way and are MiniMax-only. Marathi and Telugu are Gemini-only. Eight of the 38 languages are served by exactly one engine, and on those eight there is no decision to make — there is only a check to run before you plan around a voice you cannot have.

    The second-order trap is that reach and roster are different numbers. There are 210 named voices in the table, spread across 31 of the 38 languages, and the spread is steep: English carries 80, Spanish 20, Korean 15, Japanese 14, Portuguese 12, and it falls away quickly from there. German and Hindi carry three each. Arabic carries two. Seven of the 38 carry no voice entries at all, and 15 more list only generic Male Voice and Female Voice placeholders rather than described archetypes — which leaves 16 languages with a roster you can actually cast from.

    So "this engine supports Tamil" and "this engine offers you a choice of Tamil voices" are separate claims. Check the second one before you promise a client a specific read. The voice-over hub is where the per-language surface lives, including the four languages with their own dedicated pages.

    Axis 2: how you direct emotion — the split that breaks scripts

    This is the axis that costs people the most time, because getting it wrong does not error. It produces audio that is technically perfect and reads your stage directions out loud.

    The seven engines fall into two camps, and the camps are not negotiable.

    Inline-tag engines — ElevenLabs, Grok, Inworld. Delivery cues go into the script itself, at the point they apply.

    • ElevenLabs takes short bracketed cues immediately before the phrase they modify: [excited], [whispers], [laughs], [sighs], [sarcastic], [curious].
    • Grok takes paired tags, and every opening tag must be closed. Wrap only the phrase the emotion covers and leave the rest of the line untagged. An unclosed tag is undefined behaviour downstream, not a soft failure.
    • Inworld puts markup at the exact word or clause it modifies, mid-sentence, rather than parked at the start of a line.

    Chatterbox also parses in-text cues, but it sits outside the per-language table this post is scoped to, so it is not one of the seven being compared on language reach here.

    Parameter engines — Cartesia, Gemini, Qwen, MiniMax. The script stays clean and the mood goes in a separate field: emotion or style_instructions on the speech tool, depending on what the model accepts.

    Putting bracketed markup into a Cartesia, Gemini or Qwen script is the single most common mistake in this family. All three treat it as literal text. Your narrator says the word "excited" out loud, in a neutral tone, and then continues with the actual line. Versely's own auto-tagging step refuses to touch these providers for exactly this reason — for a parameter-based engine it hands the text back completely untouched rather than corrupting it.

    One discipline applies to all three tag engines: tag sparingly. A script with a cue on every sentence reads as noisier and less controllable to the model than three well-placed ones. The audio tags glossary entry covers the concept, and the prompting hub carries the per-family guidance.

    Axis 3: cloning, which is narrower than speech

    Cloning coverage is a smaller map than TTS coverage, and planning a multilingual campaign around a cloned brand voice without checking it is how a launch slips.

    In the current table, Cartesia records cloning support on 15 languages, Inworld on 15, and Qwen's voice-design path on 10. Between them that is 17 of the 38 languages with at least one cloning route. The other 21 are speech-capable and not clone-capable through those paths.

    Cartesia is the engine with the deepest voice pipeline behind it. clone_voice_from_audio turns an audio sample into a reusable voice_id that works with both the speech tool and the change_voice speech-to-speech conversion tool, and Cartesia fine-tunes can be tracked through to the voices they produce. None of that chain exists on the other six engines. If the plan involves one voice identity reused across many future scripts and re-voiced takes, that is a Cartesia argument before it is a quality argument.

    Axis 4: the one structural feature each engine has and the others do not

    Past reach and tags, two engines carry a capability that has no equivalent elsewhere.

    Gemini has multi-speaker mode. generate_multi_speaker_speech gives each speaker alias its own distinct voice and speaks them in the order their lines appear in the text. Write the script as labelled speaker turns and one call produces the whole exchange. For a two-voice dialogue, an interview segment or a podcast-style clip, this replaces a chain of separate single-voice generations stitched together afterwards, and it keeps the turn timing coherent because the model saw the whole conversation. The multi-voice dialogue capability is the direct route to it.

    Qwen has voice design. Instead of picking a voice_id off a roster, you describe the voice in words in a separate prompt field — the tool's own example is "a gravelly older narrator with a slow, confident cadence" — while the text field carries only the words to be spoken. Sampling controls sit alongside it. The catch worth knowing up front: this produces one clip, not a reusable voice identity. Re-running the same description gives you something similar, not something identical. Treat it as a character voice for a one-off, not as a brand voice you will still be using in six months. Designing a custom voice covers the mechanics.

    Keep the two fields apart on that one. The voice description is not part of the script — swap them and the model reads its own character brief aloud.

    Defaults per use case

    Run the gates in order, then take the default.

    Job Default engine Why
    English brand narration with emotional beats ElevenLabs Deep English roster, inline tags for per-phrase delivery shifts
    Two-voice dialogue or podcast clip Gemini Multi-speaker mode in one call, coherent turn timing
    A reusable cloned brand voice Cartesia The only engine with clone plus fine-tune plus voice conversion in one chain
    Wide-language localization sweep MiniMax or ElevenLabs The only two engines past 30 languages
    One-off character voice with no roster match Qwen voice design Describe the voice instead of picking one
    Conversational UGC read with mid-sentence shifts Inworld Markup lands at the clause, not the line start
    A phrase-scoped emotion inside neutral narration Grok Paired tags scope the emotion precisely

    None of these survive a language gate. If the target language is one of the eight served by a single engine, that engine is the answer regardless of what the table says.

    Two practical notes before you commit. Sample on the device the audience will actually use — a voice that reads as warm on studio monitors can read as muddy on a phone speaker, and short-form audio is almost entirely phone-speaker audio. And budget by script length rather than by engine, since what a voiceover for a script costs is driven by the words, not by which of the seven produced them. For the surface itself, text to speech is the entry point and the glossary definition covers what the category does and does not include.

    FAQ

    If MiniMax covers the most languages, why is it not always the default?

    Because coverage is a gate, not a quality signal, and MiniMax sits in the parameter camp for emotion. If the script needs delivery to shift on a specific phrase rather than across the whole read, an inline-tag engine gives you finer control. Coverage decides whether an engine is eligible; the other three axes decide whether it is right.

    What happens if I put ElevenLabs-style tags into a Cartesia script?

    They get spoken. Cartesia, Gemini and Qwen read bracketed markup as literal text with no special handling, so the output contains your stage directions read aloud in a neutral tone. Move the mood into the separate emotion or style-instructions field and leave the script clean.

    Can I clone a voice and use it in a language the engine does not cover?

    No. The cloning route and the language coverage are separate constraints and both have to pass. Seventeen of the 38 languages currently carry a cloning flag on at least one engine; the rest can be spoken but not cloned through those paths, so a multilingual campaign built on one cloned identity needs its language list checked against the cloning map first.

    Does the engine choice change what a generation costs?

    Different models carry different credit costs, and the app shows the cost for your exact settings before you confirm the generation. Pick on the four axes first — a cheaper engine that cannot speak your target language or take your delivery direction is not actually cheaper.