The still fixes everything the model would otherwise guess — the face, the product, the palette, the framing at the moment the clip begins. That is the entire reason to use it. Two clips built from the same still look like they came from the same shoot, because they did.
The prompt changes job. In text-to-video it describes the picture; here the picture already exists, so the prompt is stage direction: what moves, in which direction, how the camera behaves, what the shot lands on. Re-describing the image tends to hurt — the model works to satisfy a description that is already true and starts drifting the frame to match your wording.
This is the busiest lane in the catalog, and a large share of those models will not run any other way. When a record is marked as requiring an image, text alone is rejected rather than silently ignored.
In practice
- The input still is frame one, so a composition mistake in the still is permanent.
- Write motion, not description: "she turns to camera and smiles" beats "a woman in a red coat".
- Generate the still with an image model first — retrying a still is far cheaper than retrying a clip.
Image-to-video models
Catalog entries that animate a still you supply. 52 of the 296 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Happy Horse 1.1 Image to Video | Alibaba | Video |
| Vidu Q3 Image to Video | Vidu | Video |
| Pixverse 5.6 Image to Video | Pixverse | Video |
| Runway Gen-4.5 | Runway | Video |
| Wan 2.7 Image to Video | Wan | Video |
| Kling Video V3 Standard Image to Video | Kling | Video |
| Hailuo 2.3 Pro | Hailuo | Video |
| Hailuo 2.3 Fast | Hailuo | Video |
Browse all 27 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Feeding a soft or heavily cropped still. The output inherits the input frame, so a low-detail source yields a low-detail clip that no amount of prompt work rescues.
Where you will run into it
- AI Video Generator — Text-to-video, image-to-video, and story-to-video in one place.
- Morning Matcha Routine — UGC-style wellness reel — a clean-girl creator walks through her actual morning matcha ritual, then cuts to a cozy animated hero shot of the finished iced latte.
- NoseFresh — Nasal Strips Ad — Warm Pixar/Disney-style problem-solution product ad — Mark is worn down by a blocked, stuffy nose: struggling at his desk, snoring through restless nights, waking up drained. One NoseFresh strip later, his airways open (with satisfying inside-the-nose before/after cutaways) and every earlier scene gets its glow-up callback. 15 clips, ~60s vertical reel with calm narrator voiceover.
Related terms
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
First-last frame
First-last frame generation takes two stills — where the clip starts and where it ends — and generates the motion that gets from one to the other.
Reference image
A reference image is a picture supplied alongside the prompt so the model can copy an identity, product or style from it, without that picture becoming a frame of the output.
Prompt
A prompt is the written instruction a generative model reads to decide what to make — the one input almost every model requires.
Text-to-image
Text-to-image renders a still picture from a written description, with no picture going in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.