One model for image, video and audio?
FLUX 3, MiniMax H3, Wan 3.0 and Gemini Omni Flash train all three modalities at once. When one model beats a best-of-breed stack, and three jobs where it loses.
Four of the most-discussed video models in circulation right now share an architectural decision that would have looked eccentric a year ago: image, video and audio trained together, in one model, rather than assembled from three specialists at inference time. Black Forest Labs' FLUX 3, announced on 23 July 2026, is explicit about it — one model jointly trained across all three, producing video with native synchronised sound up to 20 seconds. MiniMax unveiled H3 at WAIC Shanghai on 31 July as an omni-modal model. Alibaba's Wan 3.0 opened public beta on 6 August. Google's Gemini Omni Flash has been on the same premise since its consumer launch on 19 May 2026.
The strategic question this raises is not "is the omni-modal approach better." It's narrower and more useful: for which of your actual jobs does one model beat four good ones, and where does the specialist stack still win outright? Those have different answers, and routing the wrong job to the wrong shape is one of the more expensive habits a video pipeline picks up.
What joint training actually buys
The concrete benefit is synchronisation, not capability. A model that generates picture and sound in the same pass produces audio that was conditioned on the frames it accompanies. A stack that generates video, then generates sound, then aligns the two produces audio that was conditioned on a description of the frames. Those are different products, and the difference shows up exactly where you'd expect: footsteps, impacts, mouth movement, anything where a 100ms offset reads as wrong.
That's what native audio means as a spec, and it's why it keeps appearing in the same sentence as "one model." FLUX 3 ships it to 20 seconds. MiniMax describes H3 as producing 2K video with native stereo audio in a single pass, taking up to nine reference images, three reference videos and three reference audio clips. Wan 3.0's beta claims 30 seconds at 1080p with audio in one pass, plus an "Omni-Reference" input that accepts documents, spreadsheets, slides, PDFs and webpages as generation input.
The second benefit is reference handling across modalities. A model trained on all three can accept a reference in any of them. On Versely, the Gemini Omni Video takes a prompt plus reference images, source clips, character IDs and audio IDs, under a combined quota where images plus twice the videos plus character IDs must stay at or under seven. That's a single call doing work that a specialist stack needs three calls and a conforming step to reproduce.
Where the boards actually put them
Two organisations rank video and they do not agree, so name the one you're citing. As of 14 August 2026, Arena's text-to-video board — 616,845 votes across 45 models — has gemini-omni-flash first at 1512 and flux-3-video second at 1494, with minimax-h3 sixth at 1453. Artificial Analysis's text-to-video board for models with audio has Wan 3.0 first at 1243, Gemini Omni Flash second at 1238 and MiniMax H3 third at 1227.
Two things follow. First, omni-modal models genuinely do occupy the top of both boards right now — this isn't a marketing category. Second, the two boards run on separate vote pools and separate scales, so an Elo rating from one is not comparable to an Elo rating from the other.
The three conditions where one model wins
Sound has to be locked to picture. If the deliverable has dialogue, foley or any sound event that a viewer can see happen, generate-then-sync is the wrong shape. Take the one-pass path.
The whole deliverable fits inside one generation's ceiling. This is the underrated one. FLUX 3 reaches 20 seconds; on Versely, Flux 3 Text to Video runs 5 to 20 seconds at 720p or 1080p across eight aspect ratios, at 9 credits a second. Minimax H3 Text to Video runs 5 to 15 seconds at 2K across six aspect ratios, at 26 credits a second. Gemini Omni Video runs 4 to 10 seconds, up to 4K, at 9 credits a second. If your spot is 15 seconds with sound, one model finishes it. If it's 90 seconds across six shots, you're assembling regardless and the consolidation argument weakens considerably.
Your references span modalities. A character who has to look the same, sound the same and move the same across a series is exactly the case joint training was built for. Feeding an image reference and an audio reference into the same call is structurally better than matching them afterwards.
The three jobs where the stack still wins
Stills where the text has to be exactly right. Packaging, a pricing table, a title card with a legal line — the boards are unambiguous here. Artificial Analysis's text-to-image leaderboard is led by image specialists: GPT Image 2 (high) at 1369, Reve 2.1 at 1321, Nano Banana 2 at 1320. None of the four omni-modal video models is competing at the top of that board. On Versely that means a dedicated still model such as GPT Image 2 Text to Image for the frame, then the video model for motion, rather than the video model doing both.
A voice you own and reuse. Native audio generates sound for the clip in front of it. It does not give you a persistent narrator identity that sounds the same in November as it did in August. If your channel has a voice, that voice is an asset with its own lifecycle, and it belongs in a dedicated voice cloning path rather than being re-improvised per generation.
Changing something that already exists. Real camera footage, an approved product photograph, a locked cut with a client's sign-off on it. Generation-first models regenerate; what you need is an edit that leaves everything else alone. That's a different model class — Gemini Omni Flash Edit handles conversational video-to-video at 13 credits a second, and image editing is scored on its own board, separately from text-to-image, for exactly this reason.
Run the comparison instead of arguing about it
The honest way to settle this for your own work is to put the same prompt through both routes and look at the outputs side by side, which is what the agent chat is for — one prompt, fanned across several named models in a single request, so the only variable is the model. Ask for the omni-modal one-pass version and the specialist stack version of the same 12-second spot, then judge them on the thing you actually ship on.
Two practical notes. Assemble in the editor rather than in your head: it's EDL-based, so the timeline is re-renderable and a preview: true pass gives you a 480p check at no credit cost, subject to a short per-user cooldown, with a single charge on the final export regardless of how many clips it contains. And remember that generations are billed in credits at rates that vary by length and resolution, so the comparison you want is quality per credit on your own brief rather than on a generic prompt someone else wrote.
One more thing worth knowing before you commit a pipeline. Availability is not uniform across these four. FLUX 3 launched with API and private weight access, with an open-weight "FLUX 3 Dev" backbone promised later. MiniMax has announced open weights for H3. Wan 3.0's weights status is contested across sources with no confirmed checkpoint published, so treat it as closed until Alibaba says otherwise. And Wan 3.0's beta reaches through Alibaba's own surfaces with a full API still described as coming. A model you cannot call is not a model you can plan around.
FAQ
Is one omni-modal model cheaper than a specialist stack?
Not automatically, and the direction depends entirely on length. A one-pass generation is one charge instead of three or four, which helps on short deliverables. But the per-second rates on frontier video models are the largest line in most projects — H3's listing bills more than twice what FLUX 3's does per second — so a long omni-modal clip can easily cost more than a short specialist clip plus a cheap still plus a voice pass. Price the actual brief, in credits, before deciding.
Does native audio replace a voiceover workflow?
For diegetic sound, largely yes. For a narrator, no. Native audio is generated per clip and has no memory of what your narrator sounded like last month, which is the whole point of a cloned voice. Most teams end up using both: native audio for the world inside the shot, a consistent voice for the line over the top of it.
Which of the four can I actually use today?
FLUX 3, MiniMax H3 and Gemini Omni Flash all have callable API surfaces, and Versely lists Flux 3 Text to Video, Minimax H3 Text to Video, Gemini Omni Video and Gemini Omni Flash Edit. Wan 3.0 is in public beta through Alibaba's own surfaces with a full API described as coming, so it is not something to build a production pipeline against yet.
How do I decide without running a bake-off every time?
Write down the three conditions above as a routing rule and apply it by default. Sound locked to picture, deliverable under the model's ceiling, references spanning modalities — if two of the three are true, take the one-pass route. If none is true, you are probably paying frontier video rates for a job a still model and a voice model do better.