AI News

    Grok Imagine's Multi-Reference Update: Lock Faces, Products, and Places

    xAI's July 31 update to Grok Imagine Video 1.5 added native 1080p, voice reference, and seven-image scene control. How reference-locking works on Versely.

    Versely Team7 min read

    Grok Imagine built its early reputation on being fast and a little unhinged — great for quick, chaotic image riffs, less trusted for anything that needed to hold a character together across more than one shot. The July 31 update changes that calculus. xAI shipped text-to-video, native 1080p output, a voice-reference system, and multi-reference scene control across up to seven images for Grok Imagine Video 1.5 in one release. The headline feature isn't the resolution bump. It's the reference system, which is a real answer to the question every AI ad and story creator eventually hits: how do I keep this specific thing the same while everything else in the scene changes?

    Multiple portrait reference photos pinned to a mood board

    What shipped on July 31

    The update bundles four capabilities into Grok Imagine Video 1.5: text-to-video generation (no starting image required), native 1080p output, a voice-reference system, and multi-reference scene control accepting up to seven images per generation. xAI's own release notes describe the reference system as supporting up to seven references per generation, each one able to anchor a different element of the scene.

    That last part — up to seven simultaneous references — is a meaningfully larger reference budget than most competing video models offer, and it's the piece that turns Grok Imagine from a single-shot generator into something closer to a scene-continuity tool.

    Each reference locks one thing, not the whole image

    The mechanism is the interesting part. Rather than treating a reference image as "make it look like this," each image in the set can be assigned to lock a specific element — a face, a product, or a location — while everything else in the prompt is free to change. TestingCatalog's coverage of the update puts it plainly: one reference image can hold a face in place, another can preserve a product, and another can define a location, independently of each other.

    In practice that means you can:

    • Keep a character consistent while swapping the scene around them.
    • Keep a location fixed while swapping which character appears in it.
    • Keep both character and location locked while changing only the action.

    This is character consistency handled at the reference layer instead of through prompt engineering alone, and it's a direct response to the single biggest complaint about AI video across every model in this category: the same "person" looking like three different people across three generated clips.

    Voice reference: the same face, the same voice, every scene

    The update pairs the visual reference system with a voice-reference feature. You provide a character image and a voice sample, and the two stay linked across scenes so the same face and the same voice recur in every generation that uses that character reference. For anyone building a recurring host, spokesperson, or narrator across a multi-scene ad or explainer, this closes a gap that used to require stitching a separately generated voiceover on top of video that had no guarantee of lip-sync or vocal consistency to begin with.

    A reference-discipline walkthrough on Versely

    Versely carries Grok Imagine Video and Grok Imagine Extend. Here's a concrete way to use reference discipline for a three-scene product spot inside the AI video generator:

    1. Reference 1 — the presenter. A clean portrait of your on-camera talent (or an AI-generated consistent face you've used before).
    2. Reference 2 — the product. A studio shot of the item, isolated, showing the label and shape clearly.
    3. Reference 3 — the location. A single still of the environment: a kitchen counter, a desk, a storefront.

    With all three references loaded, write scene-specific prompts that vary only the action: "The presenter picks up the product from the counter and turns it toward camera" for scene one, "The presenter sets the product down and gestures toward it while speaking" for scene two, "Close-up push-in on the product as the presenter's hand rests beside it" for scene three. The face, the product, and the counter should hold steady across all three generations because each is pinned to its own reference rather than re-described from scratch every time — the actual discipline is one reference per element, not one reference per scene.

    This is a meaningfully different workflow than the old approach of writing an increasingly long, increasingly fragile single prompt and hoping the model remembers what you said three sentences ago about the label design.

    When 30-second durations make Grok the budget long-take pick

    Grok Imagine Video supports durations up to 30 seconds in a single generation on Versely, running on a resolution-based per-second billing model — 5 credits per second at SD and 7 credits per second at HD/4K, so a full 30-second HD take lands at 210 credits. That's a genuinely long single take by current AI video standards, where a lot of competing models still cap out well under 15 seconds per generation.

    If you need to go beyond what a single generation covers, Grok Imagine Extend continues an existing Grok Imagine clip from a chosen timestamp in 2-to-10-second increments, at 6 credits per second SD and 8 per second HD/4K — the same reference-and-continuity logic applies, since you're extending footage that already has the character and voice reference baked in rather than starting a new generation cold. For the general craft of stitching generations into a longer sequence, see our guide on extending video length.

    FAQ

    How many reference images can I use in one Grok Imagine Video generation? Up to seven, each capable of anchoring a different element — face, product, location, or other visual anchor — independently of the others.

    Does the voice reference require the same image as the face reference? The voice-reference system pairs a character image with a voice sample so the same face and voice recur together across scenes generated with that character reference.

    What resolution does Grok Imagine Video output now? The July 31 update added native 1080p alongside existing lower tiers. On Versely, Grok Imagine Video is priced by resolution tier, with 1080p available at the HD/4K rate.

    Is Grok Imagine Video good for long single takes? Durations run up to 30 seconds per generation on Versely, which is long relative to most competing video models — useful for a single continuous shot rather than cutting between several short generations.

    How is this different from just writing a longer prompt? A long text prompt still asks the model to remember and re-derive every visual detail from scratch each generation. Reference-locking assigns specific images to specific elements up front, which is a more reliable mechanism for consistency than prompt length alone.

    Closing takeaway

    The multi-reference update is xAI treating Grok Imagine less like a novelty generator and more like a production tool — one where a face, a product, and a location can each be pinned down independently while the rest of the scene stays flexible. Pair that reference discipline with the 30-second duration ceiling and Grok Imagine Video becomes a legitimate option for a budget-conscious long take, not just a quick still-image riff engine. Start with three tightly cropped references — one per element you actually need to hold constant — before you touch the prompt itself.