Lunchboxes: first-last closed shell to packed tray if you shot both
Two stills. Interpolate. Do not let the model invent the sandwich.
Two stills. Interpolate. Do not let the model invent the sandwich.
A lunchbox brand is selling a shell, a latch, and a tray layout. The filling is whatever you packed that morning. Text-to-video will invent a cousin box and a meal nobody packed. Image-to-video from one closed still will guess the interior. First-last is the job only if you already photographed both states of this box.
Closed shell and packed tray are both plates you shot
Frame A is the closed shell: latch seated, lid down, brand mark readable, same bench, same light. Frame B is the packed tray: compartments filled with food you actually put there, same box, same standing spot. Those two files are the brief. A paragraph that says "cinematic lunchbox opens to reveal a healthy sandwich and fruit" is not a brief. It is a grocery list for a model that does not have your SKU.
First-last frame turns an open prompt into a bounded landing. Ordinary image-to-video knows where it begins and improvises the rest. Give it a destination and it has to land there. Packed-tray is a destination. If you never attach that plate, there is nothing honest to land on, and the sandwich will be invented.
The two frames have to be connectable in the time available. Same lunchbox, same lens height, same lighting family. A hero still of the closed shell on marble and a phone snap of a different tray on a picnic blanket is a warp, not a pack. Lock the tripod or lock your feet. Then interpolate.
Do not generate frame B. A "packed" still the model painted is a different meal and often a different hinge. Clients who own the box will notice the latch before they notice your logo.
Interpolation is the product turn, not a packing film
Create a first-and-last-frame transition video is the named agent job: two photos in, one clip that opens on the first and finishes on the second. The agent passes them as first and last frame and picks a first-last-capable model. Cost is priced per model, shown before you confirm. This is not the plain "turn one photo into a video" capability. Missing the packed tray is a different job.
Prompt the travel, not the objects. The stills already contain the shell and the food. Name the mechanism: lid opens, camera locked, no new box, no extra garnish. If the endpoints share geometry, the morph reads as a pack. If they do not, you get a mushy crossfade you could have made in an editor without a video bill.
Inspect the midpoint. Ends are held by construction. The middle is where a generator cheats — a second latch, a sandwich that swims, a tray that grows a well you do not sell. If the middle is wrong, reroll the interpolation. Do not regenerate both stills "to match the motion." The stills are the contract. The motion is the cheap middle.
A loop (the same closed still in both slots) is ambient wiggle on one plate. It is not a pack. This offer needs two states.
VEO First Last Frame holds both ends in a short take
VEO First Last Frame is the catalog row named in this job: image-to-video between a first frame and a last frame, 4 / 6 / 8 seconds, official 4K tier, 40 credits, native audio on. Audio is on — write the mix (a latch click, a quiet bench) or write silence. A silent brief still gets a soundtrack you did not choose. Score later if the bed fights the SKU.
Duration should match the morph, not a Reels quota. Eight seconds of a lid opening is a long miss if the stills are already honest. Four or six is usually enough travel. Do not ask the row for thirty seconds. That duration is not on this slug.
Do not describe the sandwich in the prompt. If the packed tray is in frame B, the food is already pixels. Language that names ingredients is how the middle grows a second filling. Language that names a different latch is how you sell a cousin SKU.
Nearby /for pages are lifestyle. This is the box.
Personal chefs book from plated dishes in a client's kitchen. That page is a tasting-menu funnel. It is not a lunchbox latch. Ecommerce brands is a general D2C surface. Do not rewrite either. The lunchbox job is closed shell to packed tray on the unit you sell.
If you only have the closed shell, wait for the packed tray or ship a still. Animating one photo and hoping the lid reveals "something healthy" is how the model invents the sandwich. That filling is not your recipe, not your allergen line, and not your SKU.
If you needed a founder line about packing with kids, that is a different generate, cut after the tray lands. The interpolation does not have to talk.
FAQ
Can I animate one photo of the closed box and hope it opens onto lunch?
You can. The ending is then the model's guess. The packed tray is the product state. Pin it as the last frame or you are not selling the pack.
What if I never shot the packed tray?
Then you have a closed-shell still, not a transformation. Pack the box. Photograph it. Do not prompt a video model to invent the filling of a real SKU. Generated grapes are a different lunch.
Do the two photos have to be shot on the same day?
They have to share geometry. Same box, same lens family, same lighting neighbourhood. A noon still into a window-lit still is a time-of-day morph you should not ask for unless that is the ad. A different hinge is a different product.
Is this the same as uploading one image to image-to-video?
No. One image is open-ended motion. Two images is a bounded landing. The agent notes that this capability needs first frame and last frame. Missing the packed tray is a different job.