Gum packs: GPT Image the flavor line, then I2V the sleeve
Legible wrapper type first. Motion only after the words lock.
Legible wrapper type first. Motion only after the words lock.
A gum ad that ships is a readable lockup on this sleeve: flavor name, piece count, the word a buyer will search. It is not a cinematic chew where the model invents a cultivar, a cousin pack, and a smear that almost says Spearmint. Lettering is a still job. The turn is a motion job. Asking a video model to typeset “Peppermint · 15 sticks” while it invents steam is how you get a misspelled wrap on a flavor that is not in the catalog.
The flavor line is the SKU
Gum demand is a name-and-count question. The frame that answers it is the actual sleeve a shopper will pick: foil or paper wrap, stick or pellet, the flavor string they already typed into search. Inventing a prettier blister does not answer “is this the 15-stick peppermint.” It answers “does this look like a movie about chewing.”
Do not generate a sleeve and present it as yours if the type is wrong. Photograph the actual pack or typeset the actual strings. The still is the contract. Motion is allowed only after the still is true. Most photoreal video models will give you a pack-shaped texture with a word-shaped smear. That is not a flavor. Keep the words in a still engine that can hold them.
Ecommerce brands already have a storefront page: PDP loops, a pack that has to survive a cart. This listing is not that page rewritten. The job here is flavor line plus sleeve.
GPT Image 1.5 letters the wrapper
GPT Image 1.5 is OpenAI’s row for conversational prompt understanding across text-to-image, image-to-image, and edits. It is strong at following long, specific instructions and rendering legible in-image text. Content type is image. Max output is 4K. Display price is 2 credits. Audio is null. There is no duration because this is not a clip.
Quote the string. "Wintergreen" and "15 sticks" are the brief, not “a premium gum pack with some lettering.” Case, alignment, how many lines, foil versus paper. Proof at 100%. A wrong character is a different SKU.
If the wrap is already on the pack in a photo you took, keep the photograph. GPT Image 1.5 is for the graphic you do not have yet, or for an edit on a locked still: one object, one instruction, the rest of the plate stays. A fresh text-to-image from the same prompt is a new sample. It will be 4K and close, and it will not be the file you approved.
Aspects include 1:1, 2:3, and 3:2. Qualities include SD, HD, and 4K. If the locked file is 4K, stay on 4K. Overlay legal lines in the editor. A flavor name is a short string this row can hold. A warning block is an editor job. Video models are a bad typesetter. Generate the lockup as a still. Then move the sleeve.
I2V the sleeve after the words lock
Image-to-video starts from the frame you already own. Frame one is the photograph or the GPT Image lockup you proofed. The prompt describes only the motion: a slow orbit, a sleeve twist that does not change the closure, a hold on the flavor line. Do not re-describe the botanicals. If you find yourself writing “whole peppermint vines in golden steam,” you have left the still and gone back to text-to-video. Text-to-video will invent a mint. The catalog does not sell that mint.
Duration is a dial on the I2V row, not a reason to switch to a cinematic talking model. A three-second sleeve turn plus a caption is an ad. An eight-second invented pantry is a mood board you will still have to caption, with a cultivar you cannot ship.
Do not prompt the model to “improve” the foil or grow a sprig that is not in the formula. Those are claims. If you “fixed” the flavor by regenerating the whole still, the video will animate the wrong pack. Edit first. Approve again. Then buy motion. Lock the 4K still at 2 credits, then pick an image-to-video row that requires that still.
Overlay the claim you can edit
Muted autoplay is how these clips get watched — in a feed, next to a checkout, with the sound off. The flavor name, the piece count, and the pack type belong on screen. Speech in the take is optional. The overlay is not.
Type the line. “15 sticks, wintergreen” is a claim with a SKU behind it. Add a text overlay burns one authored line, whole clip, top / center / bottom. It does not transcribe. If the legal line will change before the formula does, keep the wrap clean and overlay the claim after picture lock.
Two generates, one SKU. Read every character on the still against the brief. Then I2V the approved sleeve. Do not re-describe the mint in the motion prompt.
FAQ
Can I typeset the flavor in the video prompt so the sleeve letters itself?
Do not. Prompted type in a video generate is a miss you cannot edit, and the model will still invent a pantry around it. Generate the lockup as a still on GPT Image 1.5, or keep a photograph. Animate the sleeve with image-to-video.
Do I still need GPT Image 1.5 if the name is already in the photo?
No. Keep the photo. GPT Image 1.5 is for type that is not in the file yet, or for an edit on a locked still. Animate the real pack with image-to-video. Do not pay a lettering model to redraw a sleeve you already shot.
Why will a video model invent a flavor?
Because text-to-video has no plate. You asked for “cool mint steam over a rustic pack,” and the model has to decide which mint. Average is not your formula. Average is a fake cultivar with a smear for a name.
Where do allergen and “not candy” lines go if they will not fit on the sleeve?
Off the generate. Burn or overlay them after the take. A flavor name is a short string GPT Image 1.5 can hold. A full warning block is an editor job, not a prompt.