Nothing carries between generations by default. Each job starts fresh, so a face described in words is re-invented every time — close enough to feel deliberate, different enough that a viewer watching three shots in a row notices someone else arrived. This is the single biggest gap between a set of impressive clips and a piece of content that reads as one production.
The reliable fixes all work by supplying pixels rather than adjectives. Reference images give the model the subject to copy. A locked first frame carries the subject exactly into a clip. Generating every shot from one master still, rather than from one master paragraph, is the cheapest discipline that works.
Products are stricter than faces. A viewer forgives a slightly different jawline and will not forgive a label with the wrong number of words on it, so packaged goods need reference support even when a character could survive on prompt alone.
In practice
- Build one canonical still of the subject first, then derive every shot from it.
- Keep wardrobe, lighting and lens language identical across shots — those read as identity too.
- Check continuity on details a viewer scans: logo, hair parting, jewellery, sleeve length.
Models that take reference images
Models with a reference slot, and how many images each will hold at once. 40 of the 331 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Nano Banana 2 | Image | |
| GPT Image 1.5 | OpenAI | Image |
| Seedream 4.5 | ByteDance | Image |
| HunyuanImage 3.0 Instruct Edit | Hunyuan | Image |
| Qwen Image Edit 2511 | Qwen | Image |
| Flux 2 Klein 9B Base | Flux | Image |
| Qwen Image Edit | Qwen | Image |
| GPT Image 1 Edit | OpenAI | Image |
Browse all 21 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Describing a person in more detail and expecting that to hold them steady. A longer description narrows the range slightly; it does not pin the identity.
Where you will run into it
- AI UGC Video Generator — UGC ads at the speed and price of a prompt.
- AI Avatar Generator — Pick a face, hand it a script, get a presenter.
- Plush Lipstick Influencer Review — Sara the influencer reviews the Plush Baby Pink Lipstick in 3 scenes: greeting, product review with features, and promo code CTA. Consistent voiceover narration throughout.
- Podcast Clip — Versely — Podcast-clip style talking-head reel — a creator in a cozy studio teases her income jump, then reveals Versely as the tool that lets her ship videos without an editor.
Related terms
Reference image
Reference image meaning: a picture supplied with the prompt so the model can copy identity, product, or style without using it as a frame.
Reference-to-video
Reference-to-video means the model builds a clip from separate images of a subject it can place anywhere. Those images are an identity, not the first frame.
Temporal consistency
Temporal consistency is how well a generated clip keeps things the same from one frame to the next — a shirt that stays the same colour, a background that stays put.
LoRA
LoRA meaning: a small add-on file that teaches a large model a character, product, or style without retraining the whole model.
Frame interpolation
Frame interpolation invents new frames between existing ones, raising a clip's frame rate or slowing it down without it becoming choppy.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free account. Works in your browser - no install needed. The same account signs in on your phone.