When to pick VEO 3.1 Reference to Video
VEO 3.1 Reference to Video is a catalog row with a specific job. Here is when it earns the credit spend and when to route elsewhere.
Every guide, comparison and workflow we’ve published on Reference To Video.
29 articles — page 1 of 2
VEO 3.1 Reference to Video is a catalog row with a specific job. Here is when it earns the credit spend and when to route elsewhere.
Reference-to-video is files, not adjectives. Six published rows: Veo 3.1 R2V (8s, 64cr, native audio), Happy Horse 1.0 R2V (3–15s, 28cr), Seedance 2.0 Fast R2V (4–15s, 25cr, native audio), Wan 2.7 R2V (2–10s, 18cr), Kling O3 Standard R2V (3–15s, 56cr, native audio), Gemini Omni Video (4–10s, 9cr, also R2V).
pig-launchers already opened the trap door. Sleeve leftover is the split clamp around this pipe: VEO 3.1 reference-to-video is 8s, 64cr, native audio, up to 3 reference images. A generated wrap with a different bolt circle is a different repair.
Seedance 2.x @-files are multimodal references in one generate. /models/*-reference-to-video is a dedicated R2V endpoint. Different APIs, same job family.
Gemini Omni Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.
Happy Horse 1.0 Reference to Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.
Kling O3 Standard Reference to Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.
Seedance 2.0 Fast Reference to Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.
Seedance 2.0 is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.
2.5 accepts far more multimodal references than 2.0. You only need that when identity and product must lock across a sequence.
VEO 3.1 Reference to Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.
Wan 2.7 Reference to Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.
Which camera moves survive as text, which get silently approximated, and what a motion-reference clip costs you in framing freedom. With a three-bucket rule.
Two people talking needs four shots, not one. The coverage set, the eyeline rule that makes reverses cut together, and the reference discipline behind both.
Single-image, three-slot and nine-slot reference models compared: which shots improve past three references, which just get slower, and a shot-type table.
MiniMax H3 takes nine reference images, three videos and three audio clips. A working slot budget for a character-plus-product series, and what to drop first.
'It doesn't look consistent' isn't one problem. It's four — identity, wardrobe, set and lighting — each with a different cause and a different fix.
A label-fidelity workflow for AI product video: canonical reference angles, why text-faithful models matter, and the check that catches drift before publish.
Grok Imagine, MiniMax H3 and Seedance 2.5 now take stacks of reference files instead of one. How to build a brand kit and match it to the right endpoint.
ByteDance's Seedance 2.5 generates 30 seconds of synced audio and video in one pass with up to 30 references. What's new, and how to fake it today.
Editors need coverage, not a hero shot. A nine-angle shot list and a generation order that keeps one subject's identity stable across every frame.
Alibaba's Wan 3.0 beta claims 30-second single takes and video generated from PDFs, decks, and spreadsheets. A skeptical look, and what Wan 2.7 does today.
Wan 2.7 prompting guide across all three modes: text-to-video scenes, image-to-video frame direction, and reference-to-video with voice-driven speech.
Image-to-video vs reference-to-video explained: one animates your exact frame, the other recasts your subject in new scenes. When to use each, simply.