Transcription Language Coverage vs Dubbing Language Coverage
Transcription covers 165 language codes. Dubbing covers 25. Six TTS engines each cover a different subset. Here is the real coverage picture, snapshot dated.
A localization plan usually starts with one number: "the platform supports 100+ languages." Then someone actually tries to ship a Hindi-dubbed, Korean-subtitled version of a video and discovers that "supports" meant three different things depending on which step of the pipeline they were standing in. Transcription, voice generation, and full dubbing are three separate capabilities with three separate coverage lists, and the number that gets marketed is almost always the widest one — transcription — quietly standing in for all three.
That gap is not a Versely-specific quirk. It's structural: listening to speech and writing it down is a fundamentally lighter task than synthesizing a natural-sounding voice in a language, which is lighter again than cloning a specific voice's delivery and resyncing it to picture. Each step adds a requirement the previous one didn't have, and each requirement narrows the list of languages that clear the bar. Publishing the real picture, with a date attached, is the only honest way to plan around it.
Transcription is the widest net, because it's the lightest job
Transcription only has to recognize speech and place it on a timeline — no voice to generate, no lip movement to match, nothing to sound natural. That's why it's the widest of the three: Versely's caption transcription pipeline accepts 165 distinct language codes, and the list goes down to the regional-variant level rather than stopping at the language — en-US and en-GB are tracked separately, as are es-ES and es-MX, and pt-BR gets its own code apart from European Portuguese. That granularity matters more than the headline number, because a transcript tuned for the wrong regional accent or spelling convention reads as sloppy even when the language itself is technically correct.
The practical read: if the only requirement is burned-in or soft subtitles in a viewer's language, coverage is close to universal. That's the step that never becomes the bottleneck in a localization plan.
Voice generation is a per-engine roster, not one number
The moment the requirement moves from "write down what was said" to "say something out loud in that language," coverage splits by engine, because each text-to-speech engine ships its own trained language roster rather than sharing one universal list. Checked against the voice-over engine roster as of a 2026-08-11 snapshot, the named-language counts look like this:
| Engine | Languages | Notes |
|---|---|---|
| ElevenLabs | 30 | The broadest named-voice roster of the group |
| Gemini | 24 | Strong breadth, competitive with ElevenLabs |
| Grok | 16 | Mid-sized roster |
| Cartesia | 15 | Mid-sized roster |
| Inworld | 15 | Mid-sized roster, supports voice cloning in-language |
| Qwen | 10 | The narrowest named list of the group |
None of these numbers is "the" language coverage of the platform — each is the coverage of one specific engine, and picking a voice model is inseparable from picking which language list you're actually working against. A brief that needs Turkish narration and defaults to Qwen without checking will hit a wall that the same brief would clear instantly on ElevenLabs or Gemini.
Dubbing is narrower again, and for a different reason
Dubbing isn't voice generation with an extra step — it's a full pipeline that has to translate the script, clone the original delivery, and resync the result to the video, and every stage in that chain has to work for a language before dubbing into it is viable. That's why the dubbing target list, at 25 languages, sits below every individual TTS roster except Qwen's: it's not that fewer languages are "supported" in the abstract, it's that dubbing has a stricter combined requirement than any single piece of the chain.
The two available engines split the coverage further still. The default engine handles audio or video, supports trimming a clip to a specific range, and works on files up to 30 minutes — the broader of the two paths. The second engine is video-only, adds lip-sync so mouth movement matches the new language, but only supports a subset of the full target list and caps out at shorter clips with no trimming. Picking the wrong engine for a target language isn't a preference call, it's a hard compatibility check to run before promising a market a lip-synced deliverable.
For lighter jobs where full voice cloning is more than the brief needs, translating a video's spoken content without cloning the original voice covers the same 25-language target list at a smaller lift — worth defaulting to whenever the ask is genuinely "make this understandable in another language" rather than "make it sound like the original speaker said it."
The MiniMax trap: don't count style buckets as languages
One number in the voice data is worth flagging on its own, because it's the easiest one to misread. MiniMax's voice catalog totals 231 entries, but that figure collapses to 148 unique voice names once duplicates are removed, and those names sort into 33 separate buckets. That's not a language count at all — it's a set of style and character buckets (a "Wise Woman," a "Deep Voice Man," a "Movie Trailer" narrator, and so on), most of which are English-language personas distinguished by tone and delivery rather than by language. Folding those 33 buckets into a language total would make MiniMax's coverage look dramatically wider than it is, and it's a mistake worth naming explicitly rather than leaving implicit in a spreadsheet: a big roster of English character voices is a genuinely useful thing, and it is not the same axis as language coverage.
Building a plan against the real picture, not the headline number
A localization plan that survives contact with production checks coverage in this order, not the reverse:
- Confirm the target languages against the narrowest step first — if full lip-synced dubbing is required, check the dubbing target list and the specific engine's subset before anything else, because that's the tightest constraint in the chain.
- Downgrade to translated-audio-only where a full clone isn't the actual requirement. It covers the same target list at less pipeline risk and fewer engine-compatibility questions.
- Pick the TTS engine by its actual language roster for the brief, not by default — a script in Turkish or Vietnamese routed to the wrong engine is a rewrite, not a tweak.
- Treat transcription as the safe layer. At 165 codes with regional variants tracked separately, captions are close to guaranteed to clear for any language the rest of the plan is built around.
- Re-check the snapshot date before committing a market to a roadmap. Engine rosters expand over time; a plan built on stale coverage data promises languages that may not have existed when the number was pulled.
Versely walkthrough: a Hindi dub and a Korean version of the same video
A concrete case that exercises the full stack: a source video needs a Hindi voice dub for one market and a Korean subtitled cut for another. Both Hindi and Korean sit inside the 25-language dubbing target list, so a full dub_video pass into Hindi is viable on either engine, and the same tool's translate_audio_only mode covers a lighter, non-cloned Korean voice track if a full clone isn't needed for that market. Captioning either output, or captioning a version with no dub at all, draws on the much wider 165-code transcription list — Hindi and Korean are both covered there with no engine-compatibility question to run at all. Checking coverage in that order — dubbing target list first, transcription last — is what keeps a two-market localization plan from stalling on a language the pipeline was never going to hit. For picking the right engine before that first check, the best text-to-speech model comparison is the place to weigh roster size against voice quality for the specific languages a brief actually needs.
Takeaway
"Multi-language support" is at minimum three separate coverage lists — transcription, per-engine voice generation, and dubbing — and they get narrower in that exact order because each step adds a requirement the last one didn't have. Plan against the tightest constraint in the chain, not the widest number in the marketing copy, and treat any coverage count without a snapshot date attached as already out of date.