Gemini Omni Video: the stack is the brief (4s, 6s, 8s, 4K, 9cr)
Gemini Omni Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.
Gemini Omni Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.
Gemini Omni Video is Google's multimodal video row: a prompt plus optional reference images, source video clips, character IDs, and audio IDs. Catalog output is 720p, 1080p, or 4K, at 4 / 6 / 8 / 10 seconds, 16:9 or 9:16. The listed cost is 9 credits. Categories are text-to-video, image-to-video, and reference-to-video. An image is not required to press generate. The job that is actually this row is the stack you attach.
The catalog does not mark native audio on the row. Treat soundtrack as an input you may pass (audio IDs), not as a guaranteed mix you can skip writing.
The stack is the brief
A text prompt on this model is a direction, not a contract. Identity lives in what you upload. If a face, a pack, or a room has to recur, that still is part of the brief. If you do not attach it, Omni will invent a cousin and you will spend the next four takes arguing with the cousin.
The quota is arithmetic, not a vibe: images + videos×2 + character_ids ≤ 7. That is the whole budget for references. A 10-second 4K clip with seven near-duplicate hero shots is how you blow the quota and still miss the product. One face, one pack, one place is a stack. Seven moods of the same bottle is a waste of slots.
Set up reusable characters and products before you treat the generate button as a casting session. Upload once. Reference by key. The model can see a kit. It cannot remember a folder on your desktop.
Four seconds is the default, ten is a reason
The duration ladder is 4s, 6s, 8s, 10s. The slider going to ten is not a compliment. A locked product turn that reads in four seconds does not get better because you bought six more seconds of the same orbit. Use 4s to prove the stack. Step to 6s or 8s when the motion actually needs the extra beats. Use 10s when the shot is a phrase, not a glance.
Resolution works the same way. 720p and 1080p are on the row. 4K is the ceiling, not the personality. A bad stack at 4K is a sharper wrong product. Approve identity at a cheaper tier, then spend 4K on the take you are keeping.
Nine credits is the listed figure. It is still nine credits you should not spend on a prompt that never named the objects you already had as stills. The AI video generator is the launcher. The stack is the work.
Not speech, not an unbounded dump
Native audio is not marked. If the hero has to speak on camera, that is a talking or lipsync row, not a hope that an audio ID invents a mouth. The cap is 7 in that weighted formula; source video counts double. Plan the kit against the formula.
If the pack label is not approved, do not animate it here. Reference-to-video copies what you give it. It will not save a label you never designed.
The rest of Google's roster is other jobs. Omni Video is the multimodal clip with a quota. The ranked map is best reference-to-video.
A list, not a slider
Write the shot as faces, products, rooms that must appear. Empty list: text-to-video, still start at 4s. Three items: attach those three. "Longer and 4K because we can" is the slider. Spend the stack.
FAQ
Do I have to upload images for Gemini Omni Video?
No. The catalog does not require an image. Optional references are the point of the row, not a gate. If nothing in the frame has to match a real object, a prompt is enough. If something does, upload it. That is the difference between text-to-video mode and a reference-to-video job.
Why not always run 10 seconds at 4K?
Because 10s and 4K are the top of this ladder, not the default. 4 / 6 / 8 / 10 seconds and 720p / 1080p / 4K are choices. A four-second 1080p take that holds the pack is cheaper to iterate than a ten-second 4K miss that still used the same 9-credit row. Lengthen after the stack is right.
Does Omni Video replace a character kit?
No. The model accepts character IDs and reference stills; it does not file them for you. Reusable characters and products is how the same face and SKU survive the next scene. Omni consumes the kit. It does not invent the kit.
Can I pass a source clip plus seven stills?
Not if that breaks the quota. Videos count double. Images + videos×2 + character_ids must stay at or under 7. One source clip already costs two of those units. Count before you upload, or the run is a quota error wearing a prompt.