Clothing prints that morph halfway through a clip
Printed graphics, stripes and logos are the first thing a video model reinvents frame to frame. Mask that region and inpaint it instead of rerolling the shot.
Play the clip at quarter speed and watch the shirt. Not the fold. The print. A stripe count that changes. A chest logo that redrafts itself every half-second. A repeating geometric that slides, then re-tiles, then becomes a different geometric. The face held. The lighting held. The garment graphic did not, and that is the tell viewers cannot name and will not forgive on a branded shot.
This is not the same failure as cloth physics. Loose fabric that reorganises its folds is a weight-and-stiffness problem, and cloth and hair motion is the fix for that. A print that morphs is a temporal consistency failure on high-frequency, high-information texture. The model is re-drawing letterforms and repeating patterns as visual soup, the same way it re-draws small signage. You do not fix it by asking for a silkier fabric. You isolate the region and stop the rest of a good take from going in the bin.
Why prints go first
A diffusion video model does not store the graphic as a decal. It re-synthesises appearance. Anything the prompt did not pin down (the exact weave, the exact letter-spacing, the exact number of stripes) is eligible to be re-decided every frame. Prints lose that vote for three structural reasons.
They are small in the latent. A chest logo occupies a tiny share of the frame. Hands, teeth and signage fail for the same scale reason, which is why those three surfaces remain the fastest way to expose generated footage. A six-word slogan across a t-shirt is signage that happens to be worn.
They are repeating structure. Stripes, plaids, monograms, polka dots, herringbone. The model has a strong prior for "a striped shirt" and a weak prior for this striped shirt. Adjacent frames land on equally valid members of that set. Played in sequence, the set crawls.
They carry identity the rest of the shot does not. Viewers forgive a slightly different fold. They do not forgive a brand mark that rearranges itself. On a product or wardrobe-led clip, the graphic is the one surface being inspected.
Busy backgrounds, long clips and extra motion make it worse, because they spend the same consistency budget on other things. A three-to-five-second locked mid-shot of a plain tee will often hold. A ten-second orbit of a busy print almost never will. That is shot design, not model luck.
Diagnose before you regenerate
Rerolling is the expensive reflex. Classify the failure first.
| What you see | What it is | First move |
|---|---|---|
| Logo or type restyles itself; rest of the shot holds | Local texture identity | Mask and repair that region |
| Stripe count / plaid scale breathes | Repeating-pattern crawl | Mask the garment panel, or simplify the wardrobe |
| Whole garment changes cut or colour | Wardrobe identity drift | Shorter clip, or a new take with a plainer garment |
| Folds shimmer but the print is stable | Cloth physics | Weight and stiffness in the prompt, not an inpaint |
| Print is already wrong on frame one | Source still / prompt | Fix the still. Video cannot lock a graphic the first frame never had |
Two practical checks.
Compare frame one to the last frame, side by side. If the mark has become a different mark, you are looking at morph, not motion blur. If the mark is identical and only the folds moved, leave it alone.
Zoom to 100 percent on the graphic, then play. Thumbnail review passes every print. The crawl is a full-size, full-speed artefact.
If frame one is already a mushy logo, stop. Inpainting a video region cannot invent a crisp brand mark the conditioning image never contained. Rebuild the still with a blank plate or a real composite of the mark, then animate. Anything that must be pixel-perfect type is still a compositing job; mid-2026 models do not spell reliably across small type on a moving chest.
Shot design that keeps the graphic out of trouble
Repair is for takes you cannot easily get again. Avoidance is cheaper.
- Prefer solids and large blocks of colour when the wardrobe is not the product. A navy crew-neck asks the model for almost nothing. A fine stripe asks it for a decision every frame.
- If the mark must appear, make it large in frame. A back-print filling a third of the shot survives better than a tiny chest embroidery. Scale is the same rule as on-screen text: if the letters are small in the latent, they will be soup.
- Keep the shot short and the camera simple. One slow move or a locked-off frame. An orbit that reveals new parts of a patterned garment is a request to invent the pattern from a new angle, continuously.
- Do not prompt the graphic in adjectives. "A vintage band tee with an intricate illustration" licenses invention. If you must generate the garment rather than composite it, describe one simple mark ("a small white circle on the left chest, no text") and treat anything more ambitious as a post composite.
- Avoid placing the print against a competing pattern in the background. Two high-frequency surfaces in one frame is how both of them boil.
When the garment is the product, those rules invert: you cannot drop the print. Plan to generate with a placeholder and composite the real artwork, or to repair one region of a take that is otherwise perfect. Do not plan to "just get a clean take" of a small, sharp logo on moving cloth. That take is the lottery ticket, not the process.
Mask the region, do not reroll the shot
When the face, blocking and lighting are keepers and only the graphic is lying, isolate it.
1. Segment the offender. Video segmentation is the selection step. SAM 3 Video Segment takes a plain object name ("the logo on the shirt", "the stripes on the sleeve") and, when words are ambiguous, click points or a box. You are naming a region, not directing a scene. "A moody cinematic jacket" is not a selector. "The chest print" is.
2. Leave a margin. A tight mask along the ink edge leaves a seam. Feather the selection and give the fill a little garment around the graphic so light and fabric can match. Too loose, and you hand the model the face or the hands. The craft is the boundary.
3. Repair only that region. This is video inpainting: regenerate inside the mask, leave everything outside untouched. Prompt the desired graphic, not the whole shot. "Solid navy cotton, no print, no text" if you are deleting. "The same white circular mark, unchanged, sharp, identical every frame" if you are locking a simple mark. Do not re-describe the actor, the room, or the camera. The source already has those.
Instruction-style editors in the catalog are the other door when the region is a named garment rather than a drawn mask. Wan 2.7 Video Edit takes a source clip plus an edit instruction and an optional reference still of the correct print. Write the change: "Keep the jacket identical to the reference. The chest logo does not change shape, colour or lettering at any point." Set audio to origin if the soundtrack must survive; the default is not a promise.
4. If the morph is a time range, retake the range. LTX 2.3 Retake edits a segment of an existing clip from a start_time rather than the whole timeline. Use it when the first two seconds hold and only the back half of the graphic unravels. Prompt the correction, not a new scene.
5. Inspect the join, then stop. Play at full speed. If the repaired panel now sits too sharp against the rest of the cloth, the mask was too tight or the prompt asked for a sticker. Widen slightly and run again. Do not upscale until the print holds; sharpening a morphing logo is how you get a sparkling morphing logo.
Reroll the whole generation when the failure is not local: the cut of the garment changed, the actor's identity drifted, or the print is wrong from frame one. A mask cannot save a take that never had the graphic.
FAQ
Can I prompt "the logo stays identical" and skip the mask?
You can write it, and on a large, simple mark in a short locked shot it sometimes holds. Treat that as a bonus, not a plan. Naming a graphic does not pin its pixels the way a mask pins a region. If the take is otherwise expensive (a specific face, a specific gesture), budget a local repair rather than a lucky reroll.
Should I generate the real brand mark in the video at all?
Usually no. Anything that must match a real trademark, a real wordmark, or a real product label is safer as a composite: generate a blank panel or a placeholder, then overlay the actual artwork in the edit. Models still treat text as texture. A morphing fake of your own logo is worse than no logo.
Why did the print hold for a second and then unravel?
Coherence degrades as the generation runs. The opening is closest to the conditioning image; the tail is increasingly conditioned on the model's own earlier frames. That is an argument for shorter clips and for retaking only the broken range, not for a longer prompt.
Is this the same fix as a bad hand?
Same family, different surface. A hand is articulated geometry; a print is high-frequency identity. Both are local, both are cheaper to inpaint than to reroll, and both get worse if you upscale first. Segment, repair, then finish. Do not throw away a good performance because the t-shirt couldn't spell.