Garbage Reference In: Input Rules for Image-to-Video
Why image-to-video generations fail before the prompt even runs: reference resolution floors, framing, motion blur, and the clean-cutout trap explained.
Every disappointing image-to-video generation gets the same autopsy: blame the model, try a different one, maybe rewrite the prompt with more adjectives. The actual cause sits one step earlier and almost never gets a second look — the still image you fed in as the starting frame. Image-to-video doesn't interpret your photo the way you do. It has to extract depth, material, lighting direction and edges from a flat grid of pixels, and if that grid doesn't contain much to extract, no amount of prompt craft downstream fixes it.
Higgsfield's own troubleshooting guide puts weak reference input on its list of root causes behind failed generations, for essentially this reason — the model can only extract as much as the source actually contains. Three specific ways a reference ends up "weak" show up constantly in practice, and each one gets its own diagnosis below: not enough resolution, the wrong angle, and — the one people reach for as a fix and get burned by — a background stripped clean when it shouldn't have been. The reroll usually wasn't a bad draw. It was a bad input wearing a bad-draw costume.
The reference does more work than the prompt
Split the job in two. The prompt tells the model what should change — camera move, action, mood, pacing. The reference tells it what should stay the same — the subject's actual geometry, texture, and how light already sits on it in frame one. A vague prompt with a strong reference usually outperforms a meticulous prompt with a weak one, because the model can only extrapolate motion from detail that's actually present in the source pixels. Where there's no detail, it has to invent something, and invented detail is exactly where anatomy warps, materials go plasticky, and identity drifts across the clip.
That reframes the debugging question. Before touching the prompt again, look at the photo the way the model has to: no context, no memory of what the object "should" look like, just the pixels in the frame.
Resolution floor: what "low-resolution" actually means here
"Low-res" isn't just a small file. It's a lack of usable information density at the scale the model needs to work at. A 4000×3000 photo that's 90% out-of-focus background and a tiny, blurry subject in the corner is functionally lower-resolution, for the model's purposes, than a tightly framed 1024×1024 shot where the subject fills the canvas. Crop-then-check is a faster diagnostic than pixel-count-then-check: crop to roughly what the model will actually treat as the subject, and look at that crop at 100%. If it's soft, noisy, or blocky there, that's the resolution the model is really working with, regardless of what the file properties say.
The fix is mechanical, not creative: run the reference through an upscaler before it ever reaches the video model. Versely's upscale an image's resolution tool wraps exactly this step, with Topaz Upscale Image and Clarity Crystal Upscaler both in the catalog specifically for this pre-flight pass — recovering edge and texture detail in a source photo so the video model has more to hold onto once motion starts pulling the image apart. This is cheap insurance: a few credits spent upscaling a marginal reference is close to free next to the cost of a full generation that has to be discarded and rerun.
Framing and angle: the extreme close-up problem
The second failure mode is geometric rather than resolution-based. A face shot from a steep low angle, a product photographed edge-on, a subject cropped so tight that context is missing — all of these starve the model of the same thing a bad-resolution photo does: usable structural information. An extreme angle doesn't just look dramatic, it hides the surfaces the model needs to reason about as the "camera" in your generated clip moves even slightly away from that exact angle. Ask for a push-in or an orbit from a reference that only shows one steep angle of the subject, and the model is guessing at geometry it was never shown.
The practical rule: shoot (or select) the reference angle you'd actually be happy to see as the first frame of the finished clip, not a stylized crop you assumed the model would "fix." If the plan calls for camera movement, a slightly wider, more front-on reference gives the model room to move without hallucinating the parts of the subject your extreme angle cut off.
Motion blur: the failure nobody blames on the photo
This one hides well because it's invisible at thumbnail size. A source frame pulled from existing footage, a handheld phone shot, or a fast-shutter miss can carry real motion blur that reads as "fine" in a chat preview and then reads as smeared, ghosted detail once the video model starts generating new frames around it — the blur gets treated as texture and propagated, not corrected. If a reference was ever a video frame grab rather than a still shot for the purpose, zoom to 100% on the sharpest edge in the frame before trusting it. Soft edges there are motion blur, not compression, and no upscaler restores information a blurred exposure never captured — it can only slightly sharpen the blur that's already baked in.
The clean-cutout trap
The counterintuitive one: a background-removed subject on a clean plate can underperform a slightly busier real-world photo. Cutting out the background strips exactly the cues — cast shadow, contact point with a surface, ambient light direction, scale reference — that a model uses to ground the subject physically in a new scene. Feed it a floating, shadowless cutout and ask for a subject "sitting on a table," and it has to invent the entire physical relationship between subject and surface from nothing, which is a harder job than preserving one it can already see.
This matters even more when the source you're cleaning is itself a video clip rather than a still — pulling a usable, isolated plate from an existing product demo or talking-head take before reusing it as a motion reference. BRIO Video Background Removal is the catalog's tool for exactly that job on video-native footage, and it's worth applying deliberately rather than reflexively: strip a background because the new scene needs a transparent subject, not because a busy background feels like something that should be "cleaned up" by default. If the plan is a grounded scene rather than a compositing job, the un-cut original is frequently the stronger reference.
A pre-flight pass before you spend the generation
Versely's image-to-video models don't expose a "reference quality score," but the checklist above is a five-minute manual version of one, and it's worth running as a chat step rather than skipping straight to generation:
"Here's my product reference photo — check if it needs upscaling before I generate a video from it, and use Topaz or Clarity if so."
Attach the image and the agent can route it through Clarity Crystal Upscaler or Topaz Upscale Image before the actual image-to-video call, so the fix happens before credits go toward a generation instead of after a bad one. The same conversation is a reasonable place to ask which current model handles your specific subject best — Versely's best image-to-video models page tracks that ranking as the field moves, since the "best model" answer changes faster than any reference-quality checklist does.
Worth being honest about the limit here: pre-flight fixes recover information that's marginal, not information that's simply absent. An upscaler sharpens soft detail; it doesn't invent a back label the camera never photographed, and a background-removal pass can't restore a shadow it just deleted. The checklist reduces how often a bad reference silently eats a generation — it doesn't turn a genuinely unusable photo into a great one.
FAQ
How do I know if my reference is too low-resolution for image-to-video?
Crop to roughly what the model will treat as the subject and view that crop at 100%. If it's soft or blocky there, the effective resolution is lower than the file size suggests. Run it through an upscaler like Topaz Upscale Image before generating rather than after a disappointing result.
Why would removing the background make my results worse?
A clean cutout deletes contact shadows, ambient light direction, and scale cues the model uses to physically ground the subject in a new scene. If the new scene needs the subject grounded rather than composited onto something else, the un-cut original often gives the model more to work with than the "cleaner" plate.
Does a sharper reference fix a bad prompt?
No — they solve different problems. A weak reference gives the model too little structural information regardless of prompt quality; a weak prompt under-specifies what should change. Fix the reference first, since prompt detail can't manufacture information the source image never contained.
What's the difference between a reference-quality problem and a model-quality problem?
Rerun the same reference on two or three different models via reference images support. If every model struggles with the same shot, the input is the bottleneck, not any single model's capability.
Before your next image-to-video run, paste the reference into chat and ask Versely to flag resolution or framing issues before generating — cheaper than finding out after the credits are spent.