Hands, Teeth, Signage: Artifacts That Still Expose AI Footage
Why hands, teeth, and signage remain the hardest surfaces for AI video and image models — and the shot design that routes around each failure.
Lighting is convincing now. Faces hold up across a whole clip. Motion reads as physical instead of floaty. And then a hand reaches for a coffee cup, or a character smiles wide enough to show teeth, or a shop sign drifts into frame in the background — and the illusion breaks in under a second, even though everything else in the shot is doing its job.
These three surfaces aren't hard for the same reason, which is why "just generate it again" fixes one and does nothing for the other two. Each fails structurally, for a different reason, and each has a shot-design answer that avoids needing a fix at all.
Hands: articulated occlusion
A hand is one of the most geometrically demanding things a model has to render: many joints, an enormous range of plausible poses, and fingers that constantly overlap and hide each other from the camera's point of view. That's articulated occlusion — a structure with many moving, self-hiding parts — and it's exactly the condition where a model trained on the statistical average of a category tends to blur toward something that's approximately hand-shaped rather than a specific, physically consistent one. Extra fingers, fused knuckles, a thumb that bends the wrong way — none of that is the model "not knowing what a hand looks like." It's the model averaging across an enormous pose space and landing on something plausible-at-a-glance instead of correct.
Shot design that routes around it:
- Keep hands out of frame, or push them to the edge where a viewer's eye isn't drawn there anyway.
- Give a hand something rigid to hold — a cup, a phone, a railing. An object anchors the geometry and constrains the pose space the model has to solve, which measurably improves results versus an empty, freely-posed hand.
- Use motion or a shallow depth of field over a hand mid-gesture. A soft or moving hand reads as intentional camera work, not a cover-up, and it removes the sharp detail where errors are most visible.
- Cut away before a held gesture completes. The riskiest frame is often the static pose at the end of a reach, not the reach itself.
Teeth: specular regularity
Teeth are a different kind of hard. They're a tightly regular grid of near-identical glossy white objects, under lighting that changes their highlights from tooth to tooth — specular regularity, in the sense that a model has to get both the repetition (evenly spaced, similar shapes) and the variation (each tooth catching light slightly differently) right at the same time, at a scale where a viewer's eye is extremely well trained to notice anything off. The common failure isn't extra teeth so much as a generically "correct-looking" smile that reads as uncanny the moment you look for individual teeth rather than the smile as a whole.
Shot design that routes around it:
- Favor closed-mouth or subtle expressions in hero shots, especially anywhere the face is large in frame.
- Shoot three-quarter angles rather than head-on. Fewer teeth are directly camera-facing, so fewer specular highlights need to land correctly at once.
- Save the big open-mouth laugh for a quick cut rather than a held beat — the same logic as the hand mid-gesture: motion and brevity hide exactly the detail a static hold exposes.
Signage: glyph fidelity
Rendering legible text inside a generated image or video is its own distinct failure mode — glyph fidelity, the ability to draw a specific, correct sequence of characters rather than something that merely resembles text at a glance. A background street sign, a product label, a storefront name: these all ask a model to do character-level precision inside a scene it's otherwise rendering probabilistically, and that combination is unreliable even in strong models.
What makes signage different from hands and teeth is that a bad result isn't just a visual tell — on some platforms it's a compliance risk independent of whether it looks right. Amazon's main-image policy, for instance, forbids text, logos, borders, and watermarks on a product's primary image outright, no exception for whether the text is legible or garbled. A generated background sign that happens to render cleanly can still get a listing pulled from search, because the rule is about the presence of text and graphics on that image, not their quality.
Shot design that routes around it:
- Keep signage out of frame, or let depth of field blur it into unreadable background texture rather than a sharp, wrong word.
- If a scene genuinely needs legible on-screen text — a product name, a specific label — don't ask the image model to invent it inside the generation. Add it afterward as a real, controlled text layer instead of baked-in glyphs the model is guessing at.
- For commerce-adjacent shots specifically, treat "no readable text anywhere in frame" as a hard constraint on the hero image, not a nice-to-have, given what's actually at stake on a platform like Amazon.
When you can't avoid the surface: the targeted-fix path
Sometimes the hand, the smile, or the sign is the point of the shot and can't be framed away. For video specifically, the efficient move is fixing just that region rather than rerolling the whole clip. Sam 3 Video Segment can isolate the specific region — a hand, a mouth, a sign — producing the kind of segmentation mask that other tools then act on. From there, Wan 2.7 Video Edit applies a prompt-guided modification to just that segment, using an optional reference image for guidance, while preserving everything else in the source clip untouched.
A realistic sequence: identify the broken region, segment it, describe the fix in a prompt targeted at that region alone, apply the edit. You're paying for a repair on one surface, not a full re-generation of a clip that was otherwise correct.
The surfaces drift across frames, not just within one
For video, all three problems get worse over time even when a single frame looks acceptable. A hand that's merely soft in frame one can visibly reshape itself by frame thirty; a sign that's blurry-plausible in one still can flicker between two different, equally wrong words a second later. That's temporal consistency failing on top of the per-frame problem — the clip isn't just wrong, it's wrong differently moment to moment, which reads as more artificial than a single static error would. It's a second reason motion and brief framing help: a moving or quickly-cut shot never holds still long enough for the frame-to-frame drift to become visible as drift, rather than just motion.
Building it into the shot list, not catching it after
The cheapest version of all of this is deciding it before generating rather than fixing it after. A shot list that already routes hands off-frame or onto held objects, keeps smiles subtle or brief, and treats legible signage as an overlay problem rather than a generation problem produces fewer bad takes in the first place — which is strictly cheaper than a segment-level repair, even a targeted one. Whatever does make it to a final cut is worth one last look through a proper pre-publish check specifically for these three surfaces, since they're exactly the details a fast skim misses and a slow, deliberate watch catches.
None of this is about avoiding AI video's limits forever — it's about knowing which three surfaces are still genuinely hard right now, so the shot design works with that fact instead of discovering it after the render. When a shot needs a specific model for a specific surface, browse the full model catalog rather than defaulting to whatever generated the rest of the scene; the model that nails your lighting isn't necessarily the one you want anywhere near a close-up of hands.