The distinction from image-to-video matters. An image-to-video input is the first frame — its framing, angle and lighting are locked in. A reference is not a frame at all: it is an identity the model may place anywhere in the shot, at any distance, from any angle, under whatever light the prompt asks for.
That is what makes it the mode for recurring subjects. Hand it a product from three angles, or a character and the set they live in, and the same thing can appear in shot after shot without you re-shooting the opening frame each time.
The ceiling is how many references a model will hold at once. That number is published per model and is a hard limit — extra images are ignored, not blended, so a five-reference brief on a two-reference model quietly drops three of them.
In practice
- References work best when they disagree usefully: different angles beat three near-identical crops.
- The prompt still owns the shot — references decide who is in it, not what happens.
- Backgrounds inside a reference image leak. Clean plates and cutouts carry less of what you did not intend.
Reference-to-video models
Catalog entries that carry a subject from reference stills into motion. 17 of the 296 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Seedance 2.0 | ByteDance | Video |
| Gemini Omni Video | Video | |
| Happy Horse 1.0 Reference to Video | Alibaba | Video |
| Wan 2.7 Reference to Video | Wan | Video |
| Seedance 2.0 Fast Reference to Video | ByteDance | Video |
| Kling O3 Standard Reference to Video | Kling | Video |
| VEO 3.1 Reference to Video | Video |
The mistake to avoid
Submitting more references than the model accepts and assuming they were all used. Check the model's reference limit before blaming the output.
Where you will run into it
- AI Video Generator — Text-to-video, image-to-video, and story-to-video in one place.
- AI UGC Video Generator — UGC ads at the speed and price of a prompt.
Related terms
Reference image
A reference image is a picture supplied alongside the prompt so the model can copy an identity, product or style from it, without that picture becoming a frame of the output.
Character consistency
Character consistency is whether the same person, mascot or product still looks like itself across separate generations.
Image-to-video
Image-to-video animates a still you supply: the picture becomes the opening frame, and the prompt describes only what happens next.
Motion control
Motion control transfers the movement in a driving video onto a different subject — the performance stays, the performer changes.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.