The still fixes everything the model would otherwise guess — the face, the product, the palette, the framing at the moment the clip begins. That is the entire reason to use it. Two clips built from the same still look like they came from the same shoot, because they did.
The prompt changes job. In text-to-video it describes the picture; here the picture already exists, so the prompt is stage direction: what moves, in which direction, how the camera behaves, what the shot lands on. Re-describing the image tends to hurt — the model works to satisfy a description that is already true and starts drifting the frame to match your wording.
This is the busiest lane in the catalog, and a large share of those models will not run any other way. When a record is marked as requiring an image, text alone is rejected rather than silently ignored.
In practice
- The input still is frame one, so a composition mistake in the still is permanent.
- Write motion, not description: "she turns to camera and smiles" beats "a woman in a red coat".
- Generate the still with an image model first — retrying a still is far cheaper than retrying a clip.
Image-to-video models
Catalog entries that animate a still you supply. 63 of the 331 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Happy Horse 1.1 Image to Video | Alibaba | Video |
| Vidu Q3 Image to Video | Vidu | Video |
| Pixverse 5.6 Image to Video | Pixverse | Video |
| Kling O3 Pro Image to Video | Kling | Video |
| Cosmos 3 Super Image to Video | NVIDIA | Video |
| Runway Gen-4.5 | Runway | Video |
| Kling Video V3 Standard Image to Video | Kling | Video |
| Kling O3 Standard Image to Video | Kling | Video |
Browse all 29 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Feeding a soft or heavily cropped still. The output inherits the input frame, so a low-detail source yields a low-detail clip that no amount of prompt work rescues.
Where you will run into it
- AI Video Generator — Text-to-video, image-to-video, and story chaining in one studio.
- Morning Matcha Routine — UGC-style wellness reel — a clean-girl creator walks through her actual morning matcha ritual, then cuts to a cozy animated hero shot of the finished iced latte.
- NoseFresh — Nasal Strips Ad — Warm Pixar/Disney-style problem-solution product ad — Mark is worn down by a blocked, stuffy nose: struggling at his desk, snoring through restless nights, waking up drained. One NoseFresh strip later, his airways open (with satisfying inside-the-nose before/after cutaways) and every earlier scene gets its glow-up callback. 15 clips, ~60s vertical reel with calm narrator voiceover.
Related terms
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
First-last frame
First-last frame generation takes two stills, where the clip starts and where it ends, and generates the motion that gets from one to the other.
Reference image
Reference image meaning: a picture supplied with the prompt so the model can copy identity, product, or style without using it as a frame.
Prompt
Prompt meaning: the written instruction a generative model reads to decide what to make. The one input almost every model needs.
Text-to-image
Text-to-image renders a still picture from a written description, with no picture going in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free account. Works in your browser - no install needed. The same account signs in on your phone.