Audio-driven video vs video-first sound design
Generating picture from an audio track or scoring picture after it exists changes what you can still fix. A rule keyed to whichever element is locked.
There are only two orders you can generate a video and its soundtrack in, and the order decides which of the two you are still allowed to change later. That is the entire trade-off, and it is usually discovered halfway through a project rather than decided at the start.
Audio-driven generation takes a finished audio track and produces picture conditioned on it. Video-first generation produces picture, then lays sound underneath in the edit. Both end in a synced video. They differ on what is expensive to fix once the client has seen it.
What each order actually locks
In audio-driven generation the audio is the input, so it is fixed by definition. Every frame was produced in response to it. Change one word of the voiceover and you are not editing a video, you are commissioning a new one: new motion, new performance, new framing. The audio-to-video entry describes the mechanism; the consequence is that picture becomes the disposable layer and audio becomes the master.
In video-first, the picture is fixed and audio is a track you can swap without touching a frame. Replacing the voiceover is an editor operation rather than a generation. So audio becomes the disposable layer and picture becomes the master.
Native-audio models are a third thing that behaves like the second. When a model writes picture and sound in the same pass, both are locked together, and neither can be changed without regenerating both. In practice that puts it on the video-first side of the ledger for planning purposes, because the recovery move is the same: keep the picture, replace the audio in the edit. See native audio for what those models are actually producing.
Head to head
| Audio-driven | Video-first | |
|---|---|---|
| Input | Finished audio track | Prompt, image or footage |
| Lip sync | Produced by the model, tied to the take | Added as a separate pass, or absent |
| Beat and motion sync | Motion responds to the track | Cuts land on beats; motion does not |
| Picture control | Low. You describe, the audio dictates timing | High. Framing, subject and camera are yours |
| Cost of an audio change | Full regeneration | Editor operation |
| Cost of a picture change | Full regeneration | Regenerate one clip, keep the rest |
| Locked element | Audio | Picture |
The line worth staring at is the fourth one. Audio-driven generation gives up framing. If the brief includes an approved hero shot, a product angle, or anything a brand team has signed off on, that alone settles it before any of the sync arguments matter.
Where audio-driven genuinely wins
Dialogue that already exists. A recorded founder VO, a cloned brand voice, a legally reviewed script that has been read and approved. The audio is not going to change, so nothing is lost by making it the master, and you get sync from the model instead of buying it as a second pass. LTX 2.3 audio-to-video lipsync is the 1080p entry for this shape. Its predecessor, LTX 2 audio-to-video, is the flat-priced one, and the catalog states its job cost as 20 credits rather than a per-second range, which makes it easy to budget across a batch.
A still plus a track. Wan 2.2 Speech to Video takes a static image and audio and produces facial expressions, body movement and camera work from it. You are not giving up picture control you had, because you only ever had one frame.
Motion that has to feel driven by the sound. When the point of the piece is that the visuals move with the track, a model conditioning on the audio produces something an editor cutting to the beat does not. This is a genuine capability difference rather than a workflow preference.
Where video-first wins
A licensed or trending track. You cannot regenerate the music, and you do not need to. Cut the picture to it. Cuts landing on beats reads as musical to an audience; continuous motion synced to a waveform mostly does not, and costs a full generation per adjustment.
Anything with a locked frame. Product, packaging, an approved key visual.
Mixed sources. Stock, footage you shot, generated clips and stills in one timeline. Audio-driven generation has no answer for a timeline; it produces one clip per call.
The recovery path here is what makes it cheap. The editor is EDL-based, so the timeline is a description you re-render rather than a file you destructively edit. Replacing a video's audio does not touch the picture, and preview: true gives you a free 480p pass to check the result, subject to a short per-user cooldown, before you pay for the export. The final export is charged once no matter how many clips are on the timeline, which is why re-cutting is cheap and re-generating is not.
The rule, keyed to what is already fixed
Work out which of the three elements is genuinely immovable, and the order follows:
- Dialogue is fixed and picture does not have to match a specific frame → audio-driven. Feed the track and the sync arrives with the generation instead of costing you a second pass.
- Music is fixed → video-first, unless the brief specifically wants motion driven by the track. Then audio-driven, and accept that you are choosing sound over framing.
- Picture is fixed → video-first, always. No exceptions worth arguing about. Score it after.
- Two of the three are fixed → the third is your only variable, and it is the one you generate. This is usually picture, which means video-first with an added lipsync pass if there is dialogue.
- Nothing is fixed → video-first. Not because it is better, but because it keeps more options open, and briefs that start with nothing fixed rarely end that way.
The mixed case, in order. Most real jobs are case 4: approved copy, approved look, nothing else settled. The sequence that survives revisions:
- Generate the picture from a prompt or a first frame, with no dialogue in the shot. Silent or ambient only.
- Generate the voiceover separately as its own asset, so it has its own version history.
- Run a lipsync pass over the finished picture rather than asking the video model to speak. Now the audio and the picture are separable assets that happen to be synced, and the voiceover can be swapped later in the editor without regenerating a frame.
- Fix problem segments in place rather than re-rolling the whole clip. LTX 2.3 Retake takes a source video plus a prompt describing the change, with optional
start_time,durationand aretake_modethat can replace audio and video together for a chosen span. This is the tool that turns "the middle two seconds are wrong" from a regeneration into a patch. - Mix, caption and export once.
One production detail that catches people at step 5: the default timeline frame rate is 25 fps. Clips generated at a different rate get conformed, and conforming is where sync that looked fine in isolation drifts by a frame or two across a long cut. Check sync on the export, not on the clip.
FAQ
Which order gives better lip sync?
Audio-driven, on a per-clip basis, because the model is generating mouth movement in response to the actual waveform rather than being asked to match one afterwards. A dedicated lipsync pass over finished footage closes most of that gap and is usually indistinguishable at normal playback speed. The reason to still prefer the pass is not quality, it is that it keeps the audio replaceable.
Can I use both in one project?
Yes, and it is common. Audio-driven for the talking segments where dialogue is locked, video-first for the B-roll and product shots where framing matters, assembled on one timeline. The mistake is switching orders inside a single shot.
How do I keep music from burying the voiceover?
Duck the music under the speech rather than lowering it flat across the whole track, and do it in the edit rather than in the generation. Flat-lowered music sounds thin in the gaps and still competes under the words. There is a full walkthrough in ducking music under a voiceover.
Does audio-driven generation cost more?
Not inherently, but it costs more over a project with revisions, because every audio change is a new generation instead of an edit. Budget it by expected revision count rather than by unit price, and check what a video with sound costs for the per-job figures before committing a batch to one order.