Guides

    6 reference-to-video rows that want a stack, not a sentence

    Reference-to-video is files, not adjectives. Six published rows: Veo 3.1 R2V (8s, 64cr, native audio), Happy Horse 1.0 R2V (3–15s, 28cr), Seedance 2.0 Fast R2V (4–15s, 25cr, native audio), Wan 2.7 R2V (2–10s, 18cr), Kling O3 Standard R2V (3–15s, 56cr, native audio), Gemini Omni Video (4–10s, 9cr, also R2V).

    Versely Team7 min read

    Reference-to-video is files, not adjectives. Six published rows: Veo 3.1 R2V (8s, 64cr, native audio), Happy Horse 1.0 R2V (3–15s, 28cr), Seedance 2.0 Fast R2V (4–15s, 25cr, native audio), Wan 2.7 R2V (2–10s, 18cr), Kling O3 Standard R2V (3–15s, 56cr, native audio), Gemini Omni Video (4–10s, 9cr, also R2V).

    Reference-to-video extracts a subject from files you attach and restages it. A sentence that says "the same red bottle, matte cap, white label" is a hope. A stack of three stills is a brief. These six rows are the published doors. Image-to-video is a different contract: it animates the frame you uploaded. If you need the bottle on a new beach, you are on R2V, not I2V. The split is in image-to-video vs reference.

    Files are the brief

    Adjectives describe a cousin. Files name the object.

    A stack is one file per thing that must not change: the face, the pack, the room. The prompt then only names the action and the camera. If you skip the files, the model re-derives the subject from language every generate, which is how a campaign gets three bottles.

    Four of these six rows will not start without an image (requires_image: true on Veo 3.1 R2V, Happy Horse 1.0 R2V, Wan 2.7 R2V, Kling O3 Standard R2V). Seedance 2.0 Fast R2V and Gemini Omni Video will run from text, which is not an invitation to skip the stack when identity is the job. Attach the files anyway. Crop so the model can see the element you mean.

    The six rows

    1. Veo 3.1 Reference to Video — 8s, 64 credits, native audio

    Veo 3.1 R2V is an eight-second cinematic clip whose subjects match up to three reference images, with native audio. Duration is only 8s. Qualities include 720p, 1080p, and 4K. Three stills is the whole budget: one face, one product, one place — or fewer. A paragraph about "cinematic lighting" will not replace the third file.

    2. Happy Horse 1.0 Reference to Video — 3–15s, 28 credits

    Happy Horse 1.0 R2V takes up to nine reference images and addresses them as character1 through character9 in the prompt. Duration is every second from 3s to 15s. Ceiling is 1080p. Credits are 28. The published audio field is unset — do not treat this as a native-audio guarantee. Name which file is which character, then write the action. Nine slots exist so a small cast can lock, not so you fill them all.

    3. Seedance 2.0 Fast Reference to Video — 4–15s, 25 credits, native audio

    Seedance 2.0 Fast R2V is the faster Seedance reference door: up to nine images, three videos, three audio inputs, 4–15s, 720p, 25 credits, native audio on. Director-level camera language is in the description; lipsync is listed. This is the stack that can include a clip whose motion you want copied, not only stills. Keep the files few and labeled. A 30-second single pass is Seedance 2.5, a different row, text-to-video, 720p, 47 credits.

    4. Wan 2.7 Reference to Video — 2–10s, 18 credits

    Wan 2.7 R2V generates from up to five reference images and/or videos, with an optional first frame and voice-timbre control. Duration is 2s through 10s. Ceiling is 1080p. Credits are 18. The published audio field is unset. Five files is a tight stack: product angles plus one environment, or a face plus a pack. Optional first-frame is how you pin the opening without switching to image-to-video. Voice timbre is a control, not a native-audio claim.

    5. Kling O3 Standard Reference to Video — 3–15s, 56 credits, native audio

    Kling O3 Standard R2V is the character-consistency door on O3 Standard: up to four reference images, 3–15s, 56 credits, native audio on. Four stills is enough for a face, a product, a wardrobe plate, and a room — or fewer. Write motion and camera; let the files hold identity. This is not Kling O3 Pro T2V at 70 credits.

    6. Gemini Omni Video — 4–10s, 9 credits, also R2V

    Gemini Omni Video is one slug with three categories: text-to-video, image-to-video, and reference-to-video. Duration is 4 / 6 / 8 / 10s. Credits are 9. Output goes to 720p, 1080p, and 4K. The R2V path accepts up to seven reference images. The published quota is images + videos*2 + character_ids <= 7. The published audio field is unset — do not call this native audio. This is the cheap multimodal stack when you need R2V without a Veo or Kling bill.

    How to stack without writing a novel

    Crop each file to the element it is supposed to lock. A wide lifestyle still that happens to contain the bottle is a weak product reference. A clean pack shot is a product reference.

    Then write a short prompt that only names what the files cannot: the verb, the camera, the light. "She picks up the pack and turns the label toward camera, slow push, window left" is enough once the pack is a file.

    Count the slots on the row you picked. Veo: three images. Happy Horse: nine, addressed in the prompt. Seedance Fast: nine images, three videos, three audio. Wan: five images and/or videos. Kling O3 Standard: four. Gemini Omni: seven on the R2V path, inside the quota. Do not send Veo a nine-image Happy Horse stack. Run the same stack on one row until identity holds, then change only the action.

    What a sentence is still for

    A sentence is for motion, camera, and light. It is also how you pick the row in the AI video generator when there is no identity to lock — establishing weather, abstract b-roll, a room with no product. That job is text-to-video, not these six.

    A sentence is not how you keep a product honest. If the deliverable has to be this pack in a new scene, attach the pack. If it has to be this person, attach the person. Reference-to-video is files plus a short action line. The files are the brief. The sentence is the shot.

    FAQ

    Is reference-to-video just image-to-video with extra uploads?

    No. Image-to-video animates the frame you gave it, background included. Reference-to-video reads the subject from the files and restages it. Upload a studio pack shot to I2V and ask for a beach, and you still have the studio. Upload it to R2V and you can have the beach.

    Which of the six has native audio?

    The published audio: true rows on this list are Veo 3.1 R2V, Seedance 2.0 Fast R2V, and Kling O3 Standard R2V. Happy Horse 1.0 R2V, Wan 2.7 R2V, and Gemini Omni Video have audio unset. Do not assume a soundtrack policy on those three.

    How many files should I actually attach?

    As many elements as must not change, up to the row's max — usually fewer. One product still plus one face is a normal brief. Filling Veo's three or Happy Horse's nine because the slot exists is how the model gets mixed signals.

    Can Gemini Omni Video replace the dedicated R2V rows?

    It can run R2V at 9 credits, 4–10s, up to 4K. It is not Veo's eight-second native-audio cinematic, not Seedance Fast's video-and-audio reference budget, and not Kling's character-consistency door. Pin a dedicated row when that row's mix, duration, or reference types are the job.