Google · Video model

    Gemini Omni Video

    Google Gemini Omni multimodal video generation. Accepts a prompt plus optional reference images, source video clips, character IDs, and audio IDs. Supports 720p / 1080p / 4k output at 4-10 seconds (16:9 or 9:16). Quota: images + videos*2 + character_ids <= 7.

    9 credits per second of outputUp to 4K

    What Gemini Omni Video is best at

    Text to videoImage to videoReference to video
    MultimodalText to videoImage to videoReference to videoCharacter refAudio driven4k

    Pricing

    Gemini Omni Video costs 9 credits per second of output on Versely. It bills per second of output, so a full job runs 36 to 90 credits depending on clip length.

    OptionPrice
    Price9 credits

    Prices are Versely credits. Your plan's credit allowance is on the pricing page.

    Specs

    Aspect ratios
    16:99:16
    Durations
    4s6s8s10s
    Qualities
    720p1080p4K
    Max resolution
    4K

    Inputs

    Starting image

    Not required — generates from a text prompt alone.

    Video input

    Not supported.

    Release

    Gemini Omni Video was released on May 19, 2026released 3 months ago.

    Compare Gemini Omni Video

    Side-by-side pricing, resolution and rankings against models buyers weigh it against.

    Other video models

    Use Gemini Omni Video inside these tools

    Frequently asked questions

    How much does Gemini Omni Video cost on Versely?+

    Gemini Omni Video costs 9 credits per second of output on Versely. A full job runs 36-90 credits depending on clip length.

    What is Gemini Omni Video best for?+

    Gemini Omni Video is a Google video model built for text to video, image to video, reference to video.

    What resolution does Gemini Omni Video output?+

    Gemini Omni Video supports output up to 4K (available qualities: 720p, 1080p, 4K).

    When was Gemini Omni Video released?+

    Gemini Omni Video was released on May 19, 2026 (released 3 months ago).

    Try Gemini Omni Video inside Versely

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.