Dub the video or regenerate in the language
Dubbing a master and regenerating in-language are different products. Lip accuracy, cultural framing, per-market cost, and the market count that decides.
Dubbing and regenerating look like two routes to the same destination: one video, several languages. They are not. Dubbing produces a translated version of a decision you already made. Regenerating produces a different decision made in a different market. The gap between those two things is small for a product demo and enormous for anything with a joke, a price, or a cultural assumption in it.
The two operations are structurally different
Dubbing takes a finished master and replaces its audio. On Versely that runs through two genuinely different engines: an audio-cloning dub that clones the source voice and generates translated speech, working on either an audio file or a video and running up to 30 minutes; and a lip-synced translation engine that re-times mouth movement in the footage itself, video-only, capped at 8 minutes, with no trimming and a narrower language list. The two engines and their hard limits sort out which one a given source file even permits, usually before quality gets a vote.
Regenerating means writing a new script in the target language and generating the video around it. The dialogue is native from the start and nothing is being retrofitted. That structural difference decides everything downstream.
| Dub the master | Regenerate in-language | |
|---|---|---|
| Mouth accuracy | Correct only on the lip-synced engine, and only within its language subset and 8-minute cap | Correct by construction |
| Speech timing | Translated speech is rarely the source's length; needs the dynamic-duration option to fit | Written to the shot, not fitted to it |
| On-screen text | Unchanged. Audio-only operation | Native |
| Cultural framing | Fixed at master time | Decided per market |
| Cost shape | One flat job per market | A full generation run per market |
| Turnaround per market | Submit and poll | Script, generate, review, assemble |
Lip accuracy is a constraint, not a quality slider
The most common planning error is assuming lip-sync is available and then discovering it is not. The lip-synced engine is video-only, will not accept a source over 8 minutes, does not support trimming a segment out of a longer file, and covers only a subset of the audio-cloning engine's languages. Any one of those rules it out before anyone forms an opinion about the mouth.
Even where it is available, it solves one thing. Lip-sync matches mouth shapes to new audio. It does not change what the face is doing, and mismatched mouths are only the most visible half of the problem: gesture timing, eyebrow emphasis and beat placement were all performed against the original script and stay where they were. On a tight close-up held for more than a couple of seconds, viewers read the mismatch even when the mouth is right. Regeneration has no such constraint, because the performance and the dialogue were produced together and there is nothing to retrofit.
Cultural framing is where dubbing quietly loses
Dubbing changes what a viewer hears. Everything they see is a decision made for the original market and left in place:
- The setting. A kitchen, a wall socket, a car on the wrong side of the road, a cup size that reads as absurd somewhere else.
- On-screen text. Titles, lower thirds, burned-in captions, packaging in shot. All of it stays in the source language until someone re-edits, which is a different job from dubbing.
- The price and the offer. Currency, unit size, available SKU and the legal shape of the claim all vary by market. A dub can say a new number. The product shot and the lower third still show the old one.
- The hook. A hook built on a domestic reference does not become funny in Japanese by being said in Japanese. It becomes a translated joke, which everyone can identify and nobody can fix in the audio pass.
- Regulated claims. A claim that is legal at home may not be legal in the target market, and translating a claim is not the same operation as re-clearing it.
None of that is a dubbing defect. It is dubbing working exactly as specified on a master that was not built to be dubbed.
Per-market cost, and the market count that decides
The two options have different cost curves, and the curves cross.
Dubbing is flat per market. A dub bills as one job per target language, and the target language is not a pricing dimension: the charge for Japanese is the charge for Spanish. Ten markets is ten jobs and nothing else multiplies. The dubbing cost breakdown works through what does and does not multiply, including that re-running a language you already dubbed counts as a new job and there is no batch discount.
Regeneration multiplies twice: once by market, and again by however many scenes each version contains. A six-scene video regenerated for four markets is twenty-four generations plus four assembly passes, before anyone reviews anything. That is what a per-market creative build costs, and the output is a genuinely different asset each time.
So the master's structure decides the crossover point. A one-shot talking head is cheap either way. A twelve-scene product film is cheap to dub and expensive to regenerate, which is exactly when teams get tempted to dub something they should not.
Which gives a ladder by market count:
- One to three priority markets: regenerate. At this count each market can carry a real brief and a native script, the generation volume is manageable, and the difference between "translated" and "made here" is visible enough to be worth the build.
- Four to ten markets: dub the master, regenerate the top one. Your largest market gets a native build, the rest get a well-executed dub, and the total is affordable. This is where most brands actually live and where the hybrid is genuinely correct rather than a compromise.
- Ten or more markets: dub everything, and cut the master for dubbability from the start. Above ten, per-market creative stops being a localization decision and becomes running ten campaigns, which is a staffing question rather than a tooling one.
That third rule has a production consequence worth acting on before the shoot rather than after. A dubbable master is a different edit. No burned-in on-screen text where a caption track will do. No idioms, puns or domestic references in the script. Currency and price off-screen, into the caption or description. Shots framed wide enough that a mouth mismatch is not the focal point. Build the master that way and dubbing stops being a compromise for markets four through ten.
Where dubbing always loses, regardless of count
Five cases where the market count does not save you and the answer is regenerate, subtitle, or don't ship:
- Burned-in on-screen text carrying meaning. Dubbing is an audio operation. If the message is in the frame, the dub does not touch it.
- The offer differs by market. Different price, different SKU, different bundle. A dub can announce a number the visuals contradict, which is worse than not localizing at all.
- The hook is culturally specific. Translate it and you get a translated joke. Regenerate the hook even if you dub the body — a native opening on a dubbed body is a legitimate and underused hybrid.
- A jurisdictional claim difference. If the claim needs re-clearing, it needs a re-write, and a re-write needs a re-generate.
- The target language is outside the dub roster. Dubbing covers 25 target languages. Captioning transcribes across 165 language codes. For most of the world's languages the decision was made before you got to it, and automatic subtitling is the only route available — which is the coverage argument made in full in choosing subtitle or dub per market.
A practical build order
- Decide dubbability at script stage, not at localization stage. If markets four through ten are in the plan, write the master to the dubbable rules above.
- Caption the source first, in its own language. That deliverable is owed regardless of what happens next.
- Regenerate the top market with a native script rather than a translation of the master's script. Translating your own script is halfway to a dub with none of the cost savings.
- Dub the remainder through the dubbing tool, choosing the engine on the constraints — source type, runtime, whether you need a segment rather than the whole file, and whether the target language is in the lip-synced engine's subset. Use the dynamic-duration option when translated speech does not fit the original timing.
- Re-caption every dubbed output against the new audio rather than reusing the source transcript. Sentence lengths move in translation and reused timings drift.
- Fix the visuals the dub could not reach. Replace on-screen text, swap the price card, re-cut the lower third. This is the step that gets skipped and the one viewers notice.
For narration-led content that needs a natural in-language read but not one specific person's cloned voice, the voiceover library is a lighter third option: native delivery without a full dub job or a full regeneration.
FAQ
Does dubbing cost more for a harder language?
No. A dub is a flat charge per job and the target language is not a pricing dimension anywhere in the catalogue, so Japanese and Spanish bill identically. What multiplies the bill is job count: each language is one, and re-running a language you already dubbed is another. The saving available is in not dubbing markets you will not actually publish to.
Can I dub a YouTube link directly?
No. The source has to be Versely-hosted media — something from a prior generation, an upload, or an output. External links are not accepted as a dub source, so the working file has to be in the workspace before the job is submitted.
Should I translate my script or write a new one for the top market?
Write a new one. A translated script inherits every structural decision made for the original market: the order of arguments, the length of the setup, which objection gets handled first. Those are exactly the things that vary by market. Paying for a full regeneration and then translating your own script into it buys the cost of regeneration and the ceiling of a dub.
What about a native hook on a dubbed body?
It is a good hybrid and it is underused. The hook carries most of the cultural load and is a small fraction of the runtime, so regenerating the first three to five seconds per market and dubbing the rest gets a disproportionate share of the benefit for a small share of the cost. It requires the master to have a clean cut point after the hook, which is a thing to plan at edit time rather than discover at localization time.