generate_video_from_image animates a still you already have
Image-to-video is one motion job. Text-to-video, first-last-frame, and lipsync are different bills.
generate_video_from_image animates a still you already have. It is not text-to-video, not a two-photo morph, not a talking avatar. Attach nothing and you are in the wrong job. Attach a face and a voice and hope it "just talks" and you have spent a motion model on a lipsync problem.
Turn a photo into a video is the named agent job. The tagline is the spec: attach a still, get it moving. The agent asks about duration, aspect ratio, and resolution, then picks an image-to-video model. Your photo is frame zero. The clip opens on that picture and animates from there.
That is a bounded move — a product reveal, a head turn, a cloud drift, a push-in. It is not "invent the scene, then also invent the starting frame."
What the job is
One image URL. A motion prompt. A model that requires that image. The image-to-video tool is the same constraint without the chat. The AI video generator is the broader launcher; this capability is the row that refuses to start from a blank canvas.
If the still is wrong, the clip is wrong. Image-to-video will not rescue a melted extra, a bad crop, or a logo the model cannot read. Clean the still, or pick a different still, before you buy seconds.
Cost is priced per model, shown before you confirm. That price is for motion on this frame. It is not a discount on a text-to-video take you wished you had prompted instead.
What it is not
Four neighbor jobs share the word "video" and none of them are this tool:
- No photo. Generate a video from text is
generate_videos. If you do not have a starting frame, do not attach a random stock still to force this path. You will lock motion to a picture you did not want. - Two photos, a morph. First-and-last-frame transition needs both ends. One image plus "and then it becomes the other thing" is not a last frame. It is a prompt the model will approximate, then drift.
- Someone else's dance on your subject. Motion transfer needs the photo and a motion-reference clip.
generate_video_from_imagewithoutvideo_urlwill not copy choreography. - A face that speaks. Talking avatar is
generate_lipsync. Mouth shapes need audio. A "talking" image-to-video prompt gives you a face that chews the air.
Using this capability as a stand-in for any of those wastes the motion credit and still leaves the real job unpaid.
The test
Do you have one still you intend to remain obviously the same picture, only moving?
Yes: this job. Name the move. Confirm duration. Generate.
No: you either need text-to-video, two frames, a driving clip, or lipsync. Pick that row before you confirm. Honesty about the input is the whole decision.
FAQ
Can I skip attaching a photo and describe the first frame instead?
Then you are not in this job. Text-to-video invents the frame. Image-to-video is contracted to your frame. Mixing those two is a common stall: a long pause, then a question you thought you had answered.
Will a longer duration make the photo "more alive"?
It gives the model more time to leave the photo. Short clips hold identity better. If you need a longer piece, stitch later — do not ask image-to-video to become a film.
Is this the same as uploading a still into a text-to-video model that "accepts references"?
Only if that model's job is image-to-video. A reference stack on a text-to-video row is a different contract. This capability exists so you do not have to guess which catalog line requires image_url.
Why not generate the still and the motion in one prompt?
Because you will pay video rates to explore a look you have not locked. Make the still on generate_images. Approve it. Then animate it here.