Nothing carries between generations by default. Each job starts fresh, so a face described in words is re-invented every time — close enough to feel deliberate, different enough that a viewer watching three shots in a row notices someone else arrived. This is the single biggest gap between a set of impressive clips and a piece of content that reads as one production.
The reliable fixes all work by supplying pixels rather than adjectives. Reference images give the model the subject to copy. A locked first frame carries the subject exactly into a clip. Generating every shot from one master still, rather than from one master paragraph, is the cheapest discipline that works.
Products are stricter than faces. A viewer forgives a slightly different jawline and will not forgive a label with the wrong number of words on it, so packaged goods need reference support even when a character could survive on prompt alone.
In practice
- Build one canonical still of the subject first, then derive every shot from it.
- Keep wardrobe, lighting and lens language identical across shots — those read as identity too.
- Check continuity on details a viewer scans: logo, hair parting, jewellery, sleeve length.
Models that take reference images
Models with a reference slot, and how many images each will hold at once. 38 of the 296 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Nano Banana 2 | Image | |
| HunyuanImage 3.0 Instruct Edit | Hunyuan | Image |
| Kling Image O1 | Kling | Image |
| Flux 2 Klein 9B Base | Flux | Image |
| Qwen Image Edit 2511 | Qwen | Image |
| Flux 2 Klein 4B Base Edit | Flux | Image |
| Flux Kontext | Flux | Image |
| Qwen Image | Qwen | Image |
Browse all 13 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Describing a person in more detail and expecting that to hold them steady. A longer description narrows the range slightly; it does not pin the identity.
Where you will run into it
- AI UGC Video Generator — UGC ads at the speed and price of a prompt.
- Plush Lipstick Influencer Review — Sara the influencer reviews the Plush Baby Pink Lipstick in 3 scenes: greeting, product review with features, and promo code CTA. Consistent voiceover narration throughout.
- Podcast Clip — Versely — Podcast-clip style talking-head reel — a creator in a cozy studio teases her income jump, then reveals Versely as the tool that lets her ship videos without an editor.
Related terms
Reference image
A reference image is a picture supplied alongside the prompt so the model can copy an identity, product or style from it, without that picture becoming a frame of the output.
Reference-to-video
Reference-to-video builds a clip around subjects supplied as separate reference images, rather than starting from one fixed opening frame.
Temporal consistency
Temporal consistency is how well a generated clip keeps things the same from one frame to the next — a shirt that stays the same colour, a background that stays put.
LoRA
A LoRA is a small add-on file that adjusts a large model's behaviour — teaching it a specific character, product or style — without retraining or replacing the model itself.
Frame interpolation
Frame interpolation invents new frames between existing ones, raising a clip's frame rate or slowing it down without it becoming choppy.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.