Lazada listings: label words are Ideogram or overlay
Do not ask Veo to typeset the SKU.
Do not ask Veo to typeset the SKU. A Lazada buyer is reading a name, a size, a pack count, and a model string that has to match the listing title and the thing in the box. Video models draw letterforms again every frame. That is how "L-size" becomes a different glyph at second four.
Typeset the words in a still model built for type, or stamp them on after the clip exists. Veo 3.1 is a clip: native audio, official 4K tier, four / six / eight second takes. It is not a font renderer.
The listing dies on a misspelled SKU
Lazada's main image and gallery are already a compliance surface: the product in the picture has to be the product. A motion clip that "almost" shows the same label is not a style. It is a drift the buyer will screenshot. Brand names, sizes, flavour names, and barcode-adjacent strings are the expensive failures. Faces can be a little off. A bottle that gained a serif did not.
The production split is simple:
- The plate can be a photograph of the real pack, or a still generated on a type-capable image model, proofread, then locked.
- The motion is image-to-video from that plate, with a motion-only prompt, or a clean plate with no type in it.
- The words the listing cannot get wrong are Ideogram V4 on the still, or a text overlay on the finished clip.
Do not split the difference by prompting Veo "exact label, spelled correctly, facing camera." You already asked a video model to typeset. It will try. It will miss a frame. One wrong frame is the one the buyer pauses on.
Typeset in Ideogram or stamp after
Ideogram V4 is a text-to-image row: six credits a call, text rendering and posters in its catalog features. Proofread it like a print vendor proof against the listing title before it becomes the canonical pack shot. Once that still is in a reference set, the model will faithfully reproduce the typo — keeping a physical SKU believable.
Run Ideogram from text-to-image. If the listing already has a manufacturer pack photo, use the photo. Generated type is for missing pack shots, not for "the real label, but cinematic."
Once the clip exists, overlay is the other honest door. add_video_captions burns one fixed line you wrote at top, center, or bottom. It does not transcribe speech and it does not re-draw the word every frame. Several lines on a beat is timed text overlays, not a second Veo take. On-screen text in generated video is the frame-count argument: overlay draws the line once.
Veo does not set type
Keep Veo for the jobs it actually owns: a short cinematic move, native sound, 4K when the plate needs it. A Lazada gallery clip of a product orbit is a still-to-motion job if the label must hold, and a Veo job only if you generated or shot a clean plate and will stamp the SKU afterward.
A clean plate means an unlabelled pack, or a photographed pack whose type is already correct and will not be rewritten. Negative-prompt text on models that expose it so a second SKU does not appear on the table. Do not pay a 4K Veo take to invent letterforms you then paint out. The listing title is the brief. The video has to survive a pause on the label.
FAQ
Can I generate the pack in Veo and overlay the SKU on top of Veo's label?
You can, and you will fight Veo's label and your overlay at the same time. Generate or shoot a clean plate. Overlay on empty packaging. Do not stack two typesetters.
Why Ideogram instead of whatever image model is leading this week?
Because this job is lettering, not a beauty still. Ideogram V4's catalog features name text rendering and posters. Proofread anyway. A leaderboard still that cannot spell the SKU is the wrong row.
Is a photographed label safer than an Ideogram label?
Yes, when you have the real pack. A photo has ground truth on the desk. A generated label has only the proofread. Use the photo when you have it. Use Ideogram when the pack does not exist yet, then treat that file as a print proof.
Can timed overlays replace a readable pack shot?
No. Overlay carries the claim, the size, the CTA. The pack still has to be the pack. Overlay is how you stop the video model from typesetting. It is not a substitute for a correct hero still.