A first-last-frame video interpolates two stills, not one prompt
generate_video_from_image in first-last-frame mode needs two photos. One still is a different job, and faking the second wastes credits.
Two photos in. One smooth transition out. That is the job. Create a first-and-last-frame transition video is one named agent capability. The tool is generate_video_from_image, the same tool that animates a single still — but this mode passes first_frame_url and last_frame_url instead of image_url. One photo is not "close enough." One photo is a different job, and inventing the second still in prose wastes credits on a clip that never had an endpoint.
First-last-frame is interpolation. The model is contracted to open on still A and land on still B. If B does not exist, there is nothing to land on.
Two stills, or you are in the wrong row
Attach a start photo and an end photo. VEO and Kling first-last-frame models take both. The agent picks a first-last-frame-capable model and generates the clip that opens on the first and finishes on the second. The output is one video, not a pair of generates glued in an editor.
generate_video_from_image is overloaded on purpose. Plain image-to-video takes one image_url and invents motion. Motion-control variants take image_url plus a driving video_url. First-last-frame takes two stills. Same tool name, three contracts. The capability page is the contract with two stills. Do not brief it as "animate this photo toward a vibe." A vibe is not a last frame.
Cost is priced per model and shown before you confirm. That number is for this interpolation. It is not a budget you can spend to also invent the missing end still inside the same call. Generate the end still first if you do not have it — that is generate an image from text or an edit of the start — then run this job.
What this job is not
One photo, invented motion. That is turn a photo into a video. Useful. Different inputs. Feeding it as first-last-frame without a last frame is how you pay interpolation prices for a pan the single-still models already do.
A driving clip copied onto a face. That is transfer a dance or motion onto my photo. You need the photo and the motion reference video. Two stills will not copy a dance.
A longer scene. Extend a video continues a clip you already have. First-last-frame does not start from a video. If your "last frame" is actually second 7 of a take, extract the still, then interpolate — or extend the take. Do not ask this job to grow a movie.
A text-to-video generate with "start here, end there" in the prompt. Language does not pin pixels. Two files pin pixels. The AI video generator will happily spend a text-to-video row on that paragraph. You will get a clip that starts somewhere and ends somewhere else, neither of which is your still.
Pin both ends before you buy the middle
The waste is almost always the missing still. People generate a hero, like it, then ask for "a transition into the product shot" without making the product shot. The model interpolates toward a guess. You keep neither the hero identity nor a usable pack turn.
Make still A. Make still B. They should share the identity you care about — same person, same pack, same room — unless the point of the clip is the morph. Then run the first-last-frame agent job once. If the middle is wrong and the ends are right, you reroll the interpolation, not the stills. If an end is wrong, fix the still with edit a photo with AI and interpolate again. Do not regenerate both stills "to match the motion." The motion is the cheap middle. The stills are the contract.
FAQ
Can I run first-last-frame with one photo and a text description of the end?
No. This capability needs two reference photos. Describe-the-end is turn a photo into a video or a text-to-video generate. Those jobs will not land on a plate you never attached.
Is this the same tool as ordinary image-to-video?
Same tool name, generate_video_from_image. Different parameters: first_frame_url + last_frame_url versus image_url. Treating them as interchangeable is how a one-still brief spends a two-still model's credits and still has no endpoint.
Should I generate the last frame with the same model that interpolates?
Generate or edit the stills on image jobs. Interpolate on a first-last-frame-capable video model. Using the video model to invent the last frame is a generate you will then try to match, which is two identities and one seam.
When is a loop the right version of this job?
When still A and still B are the same framing — or near enough that the clip can cycle. Identical endpoints are a loop contract. Different rooms, different wardrobe, different pack: that is a transition, and it should not be asked to loop.