Which video models speak more than English
Native audio is a capability flag, not a language list. A one-generation test for any model, and the rule for generating natively versus dubbing after.
Almost every video model shipped in the last year advertises native audio, and a good number of them advertise dialogue specifically. Almost none of them publish a language list. That gap is the whole problem: "generates synchronised speech" is a capability flag, and a flag tells you the feature exists, not which languages it exists in.
Across the video catalog, exactly one entry enumerates its languages in its own description. Happy Horse 1.0 image-to-video names them: English, Mandarin, Cantonese, Japanese, Korean, German and French. Its successor, Happy Horse 1.1, says "multilingual lip-sync" and stops there. Everything else with native audio, Veo 3.1, FLUX.3, Seedance 2.0, the Sora 2 line, declares the audio capability and says nothing about coverage at all.
That is not a documentation complaint. It is the reason you have to test rather than read.
Three things that degrade, and they degrade separately
When people say a model is "worse in Spanish" they are usually collapsing three different failures that have three different fixes.
Coverage. Whether the model produces the language at all. The failure mode is not silence, it is approximation: something with the right prosody and the wrong phonemes, which sounds like the language from the next room and like nonsense up close. It is also unstable across a clip, so a model can start in the target language and drift back toward English by the end.
Accent control. Whether you can pick a regional variant. This is where video models are furthest behind the speech side, and the contrast is stark. The speech engines expose variants as explicit codes: the Grok roster carries separate entries for Mexican and European Spanish, Brazilian and European Portuguese, and three Arabic variants; the Gemini roster splits English into US, UK and Indian. On the video side you get a prompt. You can ask for a Mexican accent; nothing in the interface guarantees you got one, and nothing lets you request it the same way twice.
Mouth shapes. Visemes. Phoneme inventories differ between languages, so a model whose speech training skewed English produces English-shaped mouths regardless of what the audio says. This is the failure that survives good audio, and it is the one your audience notices without being able to name. It is also the one that gets worse the closer the shot.
Fix them in the wrong order and you waste generations. Coverage is a model choice. Accent is usually a dubbing choice. Mouth shapes are a lipsync choice.
The one-generation language test
Do not run a suite. One generation per candidate answers the question, provided the script is chosen to be diagnostic rather than representative.
Write eight to twelve seconds of dialogue in the target language containing sounds English does not have. Rolled or uvular r, nasal vowels, tones, retroflex consonants, whichever apply. Put one of them in the first three words so a failure is visible immediately, and make the last sentence the longest, because drift shows up at the end.
A workable prompt shape:
Medium close-up, single speaker facing camera, neutral indoor
background, soft key light. The speaker says, in [LANGUAGE]:
"[8-12 seconds of dialogue containing the target language's
distinctive sounds, longest sentence last]"
Natural delivery, minimal head movement, no camera motion.
Minimal head movement and no camera motion are load-bearing. They remove the two things that hide bad mouth shapes.
Then score five things, in this order, and stop at the first failure:
- Is it the language? Play it to someone who speaks it. Not "is the accent good", just "is this the language". A surprising number of clips fail here.
- Does it stay the language? Listen to the last two seconds specifically.
- Do the mouths match? Watch at quarter speed with sound off. You are looking for lip closures on the consonants that require them, which is the cheapest tell there is.
- Is the voice stable? One speaker should sound like one speaker for the whole clip.
- Is the accent the one you asked for? Last, because it only matters if the first four passed.
Run the identical prompt three times before concluding anything. Single generations are not evidence about a model, they are evidence about a seed.
Generate natively, or generate in English and dub
Once you have test results, the decision is close to mechanical.
| Situation | Route |
|---|---|
| Model passes checks 1 to 4 in your target language | Generate natively |
| Model fails coverage or drifts | English master, then dub |
| Accent is the only failure | English master, then dub |
| One master, many languages | English master, then dub, always |
| Language not in the dubbing roster | Generate natively or find a different route |
The fourth row is the one that catches teams out. Even when a model handles four of your five markets natively, generating five separate native masters means five different performances, five different framings and five different sets of notes from five different regional teams. One master plus dubbing gives every market the same picture and the same edit, which is usually what the brief actually wanted.
What dubbing gives you that native generation does not
Dubbing runs on a finished video, which means the picture is already approved before the language work starts. Two engines are available and they are not interchangeable:
- The audio-cloning engine is the default. It works on audio or video, handles up to 30 minutes, and supports trimming a section via start and end times. It clones the source voice into the target language, so the same person appears to speak it.
- The lip-synced engine is video-only, capped at 8 minutes, does not support trimming, and covers a smaller set of languages. What you get for those constraints is mouth movement matched to the new audio rather than the old.
Pick the first for anything long, anything audio-only, and anything where you only need a section. Pick the second when the speaker is on camera in close-up and the mouths would give it away. The app's roster carries 25 languages as dubbing targets, which is wider than any single video model's demonstrated dialogue coverage, and it is the practical reason the dub route wins more often than it looks like it should. AI dubbing covers the terminology, and the agent will run it directly if you ask it to dub a video into another language.
One constraint worth knowing before you plan around it: the source has to be media already hosted on the platform, from a prior generation, upload or output. External links are not accepted, so a YouTube URL is not a valid input.
What to actually do this week
If you are shipping in more than one language, spend an hour like this. Pick your two most important non-English markets. Run the one-generation test on your two default video models in both languages, three seeds each, so twelve clips. Score them on the five checks. You will almost certainly find that your default model is fine in one of those languages and not the other, which is a more useful result than any spec sheet would have given you, and it tells you exactly which market needs the dub route.
Keep the twelve clips and the date. Model versions move, and the only thing that makes a re-test cheap is having the previous test still on disk.
FAQ
Why do so few models publish a language list?
Because coverage is emergent rather than specified. These models learned speech from whatever was in the training data, so the honest answer to "which languages does it speak" is a distribution rather than a list, and distributions do not fit on a spec sheet. The one Happy Horse 1.0 entry that does publish a list is telling you which languages it considers supported, which is a stronger claim than the flag on its own.
Is a native-audio clip better than a dubbed one?
For a single short clip in a language the model genuinely handles, yes, because sound and picture came out of one pass and nothing had to be matched afterwards. Across a campaign in five languages, no, because the dubbed version keeps every market on the same approved picture. The comparison people make is clip against clip; the decision they are actually making is campaign against campaign.
Can I fix accent with prompting alone?
Sometimes, and never reliably. You can bias a model toward an accent with a prompt, and you will get it some fraction of the time, which is fine for one hero clip and useless for a batch of forty. If the accent is a requirement rather than a preference, move it to the audio layer, where regional variants are explicit selections. The Spanish voice-over page shows what that looks like when it is a setting rather than a hope.
Does the target language change what the video model costs?
No. Video is billed on duration, resolution and whether an audio pass runs, not on which language the audio is in. What changes cost is the retake count, and that is exactly what the one-generation test is for: finding out where your retake rate triples before you commit a campaign to that model. Best model with audio is the shortlist to run the test against.