AI News

    There is no Veo 4: Google's real video lineup

    There is no Veo 4. Google's video work runs across Veo 3.1, Gemini Omni Flash, Flow and Google Vids. Here is which of those you can actually name in a brief.

    Versely Team8 min read

    Veo 4 does not exist. Veo 3.1 is still the current version of the Veo line, and has been all year. That should be a boring fact, except that "Veo 4" keeps turning up in briefs, agency decks and pitch documents, usually written by someone who read a headline about a Google video feature and assumed a version bump had caused it.

    The assumption is understandable and wrong in a specific way. Google has been shipping a great deal of video capability in 2026, but it has been shipping it across four different things that get reported under the same heading: a model line, a separate omni-modal model, and two application surfaces. Version numbers advance on the first two. Features land on the last two. A brief that names the wrong one produces a procurement conversation that goes nowhere.

    The four things people mean by "Google video"

    Name What it actually is Version state
    Veo 3.1 The current version of Google's dedicated video model line Veo 3.1. There is no Veo 4
    Gemini Omni Flash A separate omni-modal model that outputs video, not a Veo successor Consumer launch 19 May 2026, API 30 June 2026
    Flow An application surface. Flow Music moved to Lyria 3.5 on 29 July 2026 Not a model, has no version you should cite
    Google Vids A Workspace product. Gemini Omni Flash landed inside it on 17 July 2026 Not a model

    Two of those four are things you can name in an API brief. The other two are places where Google's models are exposed to end users, and naming them in a technical brief is like specifying "Photoshop" when you meant a file format.

    Veo 3.1 is the Veo line, and the line got shorter this summer

    Veo 3.1 remains the version. What changed in 2026 was not a new number at the top but retirements at the bottom: Veo 2.0 and Veo 3.0 came off the Gemini API on 30 June, and the Imagen 4.0 family was retired from the same API on 17 August, which the model cleanup post covers in detail. If you have been reading that as "the line is moving fast, so there must be a 4 by now," the movement you saw was subtraction.

    On Arena's text-to-video leaderboard as of 14 August 2026, veo-3.1-audio sits at ninth and tenth with ratings of 1364 and 1363, on a board carrying 616,845 votes across 45 models. That is a solid placing for a model that has been shipping since last year, and it is also a useful reality check on the assumption that a top Google video result must be Veo.

    The Veo line as it exists on Versely is four separate jobs rather than one endpoint, which is worth knowing when you write the brief:

    • VEO 3.1 for text-to-video, at 4, 6 or 8 seconds, 720p through 4K, in 16:9 or 9:16, billed at 20 credits a second with the audio pass off and 40 with it on.
    • VEO 3.1 Reference to Video when you are driving the generation from reference material rather than a description.
    • VEO First Last Frame when you have both ends of a shot and need the middle.
    • VEO 3.1 Extend Video to push an existing clip forward in fixed seven-second hops, at 720p or 1080p.

    A brief that says "Veo" without saying which of those has not specified anything actionable, because they take different inputs and produce different cost shapes.

    Gemini Omni Flash is not a Veo successor

    This is the substitution most people are unknowingly making when they write "Veo 4." Gemini Omni Flash is a different model on a different line, built as omni-modal rather than as a dedicated video model. It launched to consumers on 19 May 2026 and reached the API on 30 June 2026.

    It is also, right now, the top entry on Arena's text-to-video board at 1512. So when someone reads that Google has the leading video model and infers a new Veo, the underlying fact is real and the inference is wrong. The leading Google video result on that board is not a Veo model at all.

    On Versely it appears as two distinct listings, which matches the two distinct jobs:

    • Gemini Omni Video generates from a prompt plus optional reference images, source clips, character IDs and audio IDs, at 4 to 10 seconds, 720p through 4K, in 16:9 or 9:16, at 9 credits a second. The reference inputs are capped by a combined quota where images plus twice the videos plus character IDs must not exceed seven.
    • Gemini Omni Flash Edit handles conversational video-to-video editing at 13 credits a second.

    If a brief says "Gemini video," the first question back should be generate or edit, because those are different models with different billing.

    Flow and Google Vids are surfaces

    Flow is an application. The thing that changed there recently is audio, not video: Flow Music moved to Lyria 3.5 on 29 July 2026, replacing Lyria 3 Pro, with vocal tracks up to three minutes, image-to-music input, and tempo and duration control. That is a real capability change and it is not a video model release, which is exactly the sort of item that gets summarised into "Google shipped new video AI" two hops downstream.

    Google Vids is a Workspace product. Gemini Omni Flash landed inside it on 17 July 2026, bringing personal avatars and text-prompt editing of real camera footage, with every generated clip carrying a SynthID watermark. Again: a surface gaining a model, not a model gaining a version.

    The distinction matters commercially because surfaces and models have different procurement paths, different regional availability and different terms. A client who asks for "Google Vids quality" is describing an outcome. A brief that specifies Gemini Omni Video at 1080p, 8 seconds, 9:16 is describing a job someone can run.

    Translating a brief into something that exists

    What the brief says What it probably means What to write instead
    Veo 4 Either the current Veo, or the model topping the video board Veo 3.1, or Gemini Omni Flash. Pick one and say which
    Gemini video Gemini Omni Flash Gemini Omni Video to generate, Gemini Omni Flash Edit to change existing footage
    Flow An app the client has seen The model and settings you actually want
    Google Vids output A Workspace deliverable The model, duration, resolution and aspect ratio
    Imagen video A category error Imagen is the image line. Name a video model
    Veo with sound The audio configuration Veo 3.1 with the audio pass on, which changes the credit rate

    Google's provider page lists what is currently callable, which is the fastest way to settle an argument about whether a name refers to something real. If the name is not there, it is either a surface, a retired version, or a model that was never released under that number.

    FAQ

    Is Veo 4 coming?

    Nothing has been announced, and the correct thing to write in a brief today is Veo 3.1. Speculating about an unannounced version in a client document is how a deliverable gets scoped against a model that may never carry that name.

    Which should I use, Veo 3.1 or Gemini Omni Flash?

    They have different shapes rather than a clear winner. Veo 3.1 gives you 4, 6 or 8 seconds with an audio pass you can switch off to halve the credit rate, plus dedicated reference, first-last-frame and extend variants. Gemini Omni Flash reaches 10 seconds, takes a wider set of reference inputs including character and audio IDs, and bills at a lower per-second rate. Run both on the same prompt before committing a series to either.

    Why do I keep seeing Google video news with no new model behind it?

    Because most of it is surface news. A model that shipped in May reaching Google Vids in July is genuinely newsworthy for Workspace users and is not a model release. The reliable filter is to ask whether the item names a version number that changed. If it does not, no model shipped.

    Does the SynthID watermark apply to everything Google generates?

    The Google Vids announcement states that every clip generated there carries SynthID. That is a statement about that surface. Whether a given model on another surface embeds a signal is a per-model question worth checking rather than assuming in either direction, and it is separate from whether a platform stamps a visible mark of its own.