Packshot on white, then image-to-video
Lock a white packshot still first. Image-to-video inherits that plate, ratio, and label. Do not invent the SKU in the clip.
Lock a clean product still first, then image-to-video. The packshot on white is the plate. Motion inherits every pixel you approved: silhouette, label, ratio, contact shadow, empty field. A text-to-video "white bottle" is a cousin SKU. A lifestyle still is a different plate. This page is the photography sequence, not a credit ranking and not a wrapper-lettering job.
White is a constraint. That is the point. A kitchen gives the sampler curtains to blow and hands to invent. A sweep leaves the object. If the object is wrong, eight seconds of 4K will not make it yours.
The white plate is the SKU
A packshot answers "is this the thing I searched." Marketplace main images are built on that answer. Amazon's seller help for product images (G1881, read 19 September 2026) still requires the MAIN slot to sit on pure white, RGB 255, 255, 255, with the product filling about 85 percent of the frame, the whole item inside the frame, no cut-off edges, no extra units, no badges. MAIN is a still. The clip you make after it is a gallery loop, a Reel, a PDP hover, an ad. Do not replace the listing hero with a generated movie and call it compliance.
That split matters for how you source the plate. Amazon wants a realistic photograph of the actual product in MAIN, not an illustration. If you sell the SKU, photograph it. If the sweep was a garage wall, edit the photo onto white. If the object is not tooled yet and the job is a concept ad, generate a packshot and treat it as concept. Do not upload a generated cousin into MAIN and hope the catalog page forgives you.
White is also the I2V constraint. A lifestyle plate (marble, linen, a window) gives the motion model a room. Rooms grow steam, palms, extra bottles, a second label in the reflection. A white sweep has almost nothing to hallucinate except the product and the void. The void is the remaining trap. A tiny bottle floating in a sea of white is empty-field motion: crawling gradients, phantom speculars, dust that was never on the glass, a second shadow that walks. Fill the frame the way the listing rule already wanted. 85 percent is photography and I2V hygiene.
Straight-on versus three-quarter is the other photography choice the clip will not forgive. Amazon MAIN is often dead-front, eye level, honest, slightly boring. That plate orbits like cardboard. The model has no side to reveal, so it invents thickness or it warps the face. For a turn, lock a three-quarter packshot on the same white sweep: face plus a sliver of the side, same light family, same fill. Keep the dead-front file for the grid. Animate the three-quarter.
Contact shadow is a still decision. A real sweep has a tight shadow where the object meets the paper, and sometimes a soft falloff as the backdrop curves. Strip both and the product reads pasted. Leave a wandering wall shadow and I2V will crawl it. Decide on purpose: flat field, or a short contact directly under the base. Prompt the camera, not a table that was not in the JPEG. If you write "sitting on a wooden table" against a white plate, you asked the sampler to grow furniture in the void.
Proof the label at 100 percent before anyone buys motion. Wrong letters in a still are a retouch. Wrong letters in a loop are a different SKU on repeat. If the type is the offer, overlay it after the take. If the type is on the pack, it has to be true in the still.
Crop the still to the clip before you generate
Image-to-video inherits the frame. Several I2V schemas do not even expose an aspect field. Vidu, Pixverse, and Wan notes in the prompting catalog are blunt: ratio is inferred from the upload. Asking for "9:16 vertical" in the motion prompt does nothing if the JPEG is a 1:1 Amazon plate. You get a square clip, or a square letterboxed into a tall export. Crop first.
Match the destination:
- 1:1. Marketplace grid, many PDPs. Generate or crop the packshot square, product still filling the frame, even margin, nothing clipped.
- 4:5. Feed posts. A 1:1 plate padded top and bottom is a postage stamp. Recrop so the bottle still occupies the tall frame.
- 9:16. Reels, Stories, TikTok. Recrop the packshot into a tall still. Keep the 85 percent fill. A tiny hero in a white tunnel is the empty-field problem again, now taller.
- 16:9. Site headers, YouTube. Same rule, sideways.
When you generate the still, pick the ratio from the model's enum. Do not type "widescreen" into the prompt and hope. Text to image is the stills door. Flux 2 Max is 6 credits, photoreal, text-to-image only. It will invent a bottle. Use it for a concept SKU, not as a substitute for a photograph of inventory. Imagen 4 is shut down on Google's API (17 August 2026). Do not open a leftover Imagen card for this job. Nano Banana 2 is Google, text-to-image plus image-to-image plus edit-image, 4 credits discounted (7 full, 50 percent off on the 19 September 2026 catalog), up to 14 references, qualities 1K / 2K / 4K, aspects include 1:1, 4:5, and 9:16. That is the row when you already have a photo and the job is "put this object on white, keep the object." An edit is not a new generate. Attach the file. Ask for the sweep. Proof the subject, not only the background.
Photograph, then edit, is the honest packshot path for a live SKU. Generate, then I2V, is the concept path. Mixing them is how a listing shows a bottle that is not in the warehouse.
Checklist before anyone opens a video row:
- Eyedropper the field. "Looks white" and RGB 255, 255, 255 are different claims. Off-white becomes a dirty loop the moment the clip sits on a true-white PDP.
- Product fill. If you can see more sweep than object, recrop.
- Whole product in frame. A clipped cap becomes a morphing cap when it moves.
- Label at 100 percent. One wrong character is a different SKU.
- Ratio matches the clip you will actually post.
- Contact shadow is a choice, not leftover junk from the garage wall.
- No badges, watermarks, or price stickers you do not want to watch crawl.
Do not upscale a muddy still and then buy motion. I2V sharpens nothing about identity. It spends the pixels you uploaded.
Motion a packshot can survive
The still already decided the product, the light, and the set. The motion prompt describes what changes. Re-describing the bottle fights the JPEG. "Premium matte serum, golden hour, marble slab, cinematic fog" is a text-to-video brief wearing an image-to-video button. The sampler will try to honor both. The marble wins often enough to ruin the plate.
Write the physical action, sized to the duration you will actually pick:
- Slow yaw, 15 to 30 degrees, camera height locked, product stays on the sweep.
- Slow push-in, no tilt, stop before the label crops.
- Specular travel along a highlight that already exists on the glass or the cap.
- A slight settle, if the still already has a contact shadow and you want the object to feel heavy.
That is most of a packshot ad. Hands, pours, steam, extra units, a kitchen, a linen napkin, a second bottle "for scale," a cap color change, a label rewrite, "make it more premium," floating, and a silhouette morph are not packshot motion. They are a new shoot. If you need a pour, lock a still of the pour (or shoot it). If you need a hand, composite a real hold. Do not ask the white void to grow a person.
Empty white plus "cinematic lighting" is how colored gels and haze appear in a field you meant to keep empty. Name the light only if you are asking it to stay: "soft studio wrap, no new gels, background stays white, no fog." Negative lists help when they name objects, not vibes: no hands, no extra bottles, no text change, no furniture, no floorboards.
Duration is a glance. A packshot turn that works at four to eight seconds usually dies at thirty. The object has nowhere to go. The model starts inventing weather in the white. Start on the short end of the row you picked. Trim in an editor if the platform cap is shorter than the block you already paid for.
Audio is a separate problem. Packshot ads are watched muted. Native audio on a white plate invents room tone, a whoosh, or a voice you did not write. If the row can emit sound, you can still prompt visual silence and overlay the claim. Price, count, and "in stock" belong in type you can edit, not in letters the video model will crawl.
If the still was the wrong ratio, stop. Recrop. A padded 1:1 inside 9:16 is not a vertical ad. It is a square with white pillars, and I2V will treat those pillars as more void to fill.
Seedance and Veo after picture lock
The generate door for the clip is the AI video generator. Filter by the input you actually have. A packshot JPEG is an image-to-video job. Pasting the same sentence into a text-to-video row throws the plate away.
Seedance 2.5 is the published family row: ByteDance, native audio, 4 to 30 seconds, 720p ceiling, min_credits 34, max_credits 557, 50 percent off, SD discounted cell 9, HD/4K discounted cell 19 (19 September 2026 models.json). The public page is filed as text-to-video. The image-to-video sibling in the catalog copies that rate card and requires the still. That sibling is not a published /models/ URL, so this page does not invent one. Bring the packshot in the studio picker. Use Seedance when the turn needs more than eight seconds and 720p is enough. A 30-second white orbit is still a screensaver. A 6-second three-quarter yaw is the usual job. Do not budget the 34-credit floor for a 30-second hero. 34 is the short SD end of the card.
VEO 3.1 is the published Veo row: Google, native audio, 4 / 6 / 8 seconds, 4K listed, min_credits 64, max_credits 256, no discount. Audio off is 16 credits a second. Audio on is 32. The 64 floor is four seconds with audio off. Use Veo when the plate needs 4K and cinematic physics on a short hold. Do not use it to invent the SKU. If the motion prompt re-stages the set, Veo will grow the marble. If the motion prompt only yaws the approved packshot, you kept the contract. Veo's published page is text-to-video with requires_image false. The still is still the job you came for. Attach it. Do not "describe the bottle instead."
Those two family pages are the published T2V doors this brief named. Other I2V rows exist in the catalog. Some have public model pages. Some do not. Quote unpublished slugs as catalog facts if you must. Do not link a /models/ path that scripts/_published-models.json does not list. This page is not a floor table. Rank the row by length, resolution, and whether you are allowed to emit sound, after the still is true.
Two generates, one SKU. Proof the white still. Crop it to the clip. Prompt only the turn. Overlay the claim you can edit. If the JPEG is locked, create a Versely account and generate from the on-screen quote.
FAQ
Can I skip the still and prompt Seedance or Veo from text?
Not if the SKU has to be yours. Text-to-video will invent a plausible bottle. Plausible is a cousin. Photograph the product, or edit a photograph onto white, then image-to-video that file. Concept work can start on text to image. Listings and ads for inventory cannot.
Why is the clip square when I asked for 9:16?
The still was square. Image-to-video inherits the upload. Recrop the packshot to 9:16 with the product still filling the tall frame, then run the motion pass. Typing "vertical" into the prompt does not recrop.
Should the motion prompt add a kitchen or marble?
No. That is a new still. White is the set you locked. A kitchen, a slab, or a hand is a different plate, locked and proofed on its own, then animated. Asking I2V to grow furniture in the void is how the SKU stops being the SKU.