Guides

    VEO 3.1 Reference to Video: the stack is the brief (8s, 64cr)

    VEO 3.1 Reference to Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.

    Versely Team4 min read

    VEO 3.1 Reference to Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far. Google's row is an 8s cinematic clip whose subjects match up to three reference images, with native audio, billed at 64 credits. That is the whole machine. The stills are the brief. The prompt is blocking, not casting.

    This is not image-to-video. Frame one is not your upload. The model recasts the people and products you showed it into a new shot. If you needed the kitchen you photographed to stay the kitchen, you are on the wrong category. The split is image-to-video vs reference.

    Three stills, one eight-second take

    The catalog caps the stack at three images. That is a feature, not a shortage. Front of the spokesperson, three-quarter of the SKU, one clean plate of the room — then stop. A fourth near-duplicate does not buy fidelity; it asks the model to average two lighting setups into a third face.

    Register the set once with reusable characters and products so clip nine pulls the same keys as clip one. Hygiene still applies: same crop family, same light, no mixing a studio packshot with a phone grab. Mixed refs are how you get a person who is in none of the photos. Reference hygiene is the spec for that.

    The prompt then names action and camera, not hair color. The stills already answered who. Re-describing the subject in text is how the stack and the sentence fight.

    Native audio is in the take

    audio is true on this row. The eight seconds come back with a soundtrack. A silent prompt does not mean a silent file. Write the room: dialogue, or none; footsteps, or none; music, or none. If you needed a mute plate to score later, say so. If you needed speech, put the line in the prompt instead of hoping Google guesses the brand voice.

    Aspect is 16:9 or 9:16. Quality settings on the row include 720p, 1080p, and 4K. Duration is not a slider. supports_durations is 8s only. There is no 4s draft of this model and no 15s hold. One length. Plan the cut for eight seconds or pick another row.

    64 credits is not a reason to max a stack you have not locked

    Requires an image. Text-only will not run. Spend the cheap stills pass on text to image until the face and the label are the ones you would print. Then spend 64 credits once. Re-rolling VEO because the product is the wrong color is how a reference model gets used as a sketchbook.

    The current ranking of the category lives at best reference-to-video. Google's other video family is on the Google provider roster. This page is only the reference row: three stills, eight seconds, native audio, 64 credits.

    If the job is "this person, this bottle, new scene," you are in the right place. If the job is "animate this exact frame," leave. The image-to-video tool is the other door.

    FAQ

    Is this the same as uploading a first frame?

    No. Reference-to-video extracts identity and restages it. Image-to-video treats your file as frame one. VEO 3.1 Reference to Video is the former: up to three stills, new cinematic eight-second clip, native audio.

    Why only eight seconds?

    That is the duration the catalog lists. There is no 5s or 10s option on this slug. Write an eight-second shot or change models. Stretching the idea past eight seconds is a second generate, not a longer slider.

    Do I have to use all three reference slots?

    No. Empty slots are better than junk. One clean hero and one product beat three mixed lighting setups. The max is three, not a quota.

    What does 64 credits buy?

    One generate on this row as listed in the catalog: an 8s clip, native audio, subjects matched to the stills you uploaded. It does not buy a pack of lengths or a silent-only mode you forgot to request. Prompt the sound you want.