Industry

    Virtual Try-On: What It Nails and Where It Lies

    Virtual try-on renders convincing fabric drape from a single photo, but a diffusion image is not a fit simulation — what retailers owe customers in disclosure.

    Versely Team8 min read

    Type "how does virtual try-on actually work" and most answers stop at "AI puts the clothes on a model." What's actually happening underneath is worth knowing before you trust the picture, because it explains both why the results look so much better than the try-on tools from three years ago, and why "it looked right on the site" is not the same claim as "it will fit."

    What's actually being generated

    Google's apparel try-on model, which is the most-documented version of this technology and the one worth using as the reference case, doesn't warp a photo of a garment onto a photo of a body — that was the old approach, and it's the reason early try-on tools produced the "misplaced folds that make garments look misshapen and unnatural" that Google's own writeup calls out. The current approach is generative: a diffusion model takes a garment photo and a photo of a person as a pair of images — not text — and two separate neural networks (U-nets) process each one while sharing information through cross-attention, so the output is a newly rendered image that has to be plausible as a photograph, not a pixel-warp of the input.

    The shopper doesn't upload their own photo, either. What you get to do is pick from a set of real models with different builds — Google's own framing is seeing "what a top looks like on the model of your choice" — and the model renders the garment onto that chosen body. That's an important detail for what comes later: the system was never claiming to show you. It's showing a photograph-plausible version of the garment on someone whose proportions you chose because they're closer to yours than a single house model would be.

    Where it earns "convincing"

    The properties a diffusion-based try-on model reproduces are specifically the ones a flat product photo can't show and a geometric warp gets wrong: drape, folds, cling, stretch, wrinkles, and shadow, generated from a single clothing image. Those six words matter more than they look — they're the visual cues a shopper's eye actually uses to judge whether a fabric moves like cotton or like something stiffer, whether it clings at the waist or hangs loose, whether the hem catches light the way real fabric does under real light. A warp-based try-on gets the garment's outline roughly right and the physics completely wrong; a generative one is, by construction, trying to produce something that reads as a photograph of fabric behaving like fabric.

    The size-and-diversity side of this is a genuine, separate advance rather than a marketing footnote: Google describes a model set spanning a wide size range with skin tones guided by the Monk Skin Tone Scale, which is the difference between "here's how it looks on the one body type we photographed" and "here's how it looks on someone whose shape is actually close to yours." For a shopping decision, that's not cosmetic — a garment that drapes cleanly on a slim frame and pulls at the seams on a fuller one is showing you two different products, visually, and a diverse model set is the only way an on-model photo tells the truth for more than one kind of customer.

    Where it lies: fit isn't in the image

    Here's the part that doesn't survive contact with a return department. A diffusion model, no matter how good it is at drape and cling, is generating an image — a plausible arrangement of pixels conditioned on a garment photo and a body photo. It is not running a physical simulation of that exact fabric under tension on that exact body. Those are different problems, and the model only solves the first one.

    Concretely, three things a convincing on-model render cannot tell you:

    • Whether the size is right for you. The model you picked is wearing a specific size in a specific cut. Your measurements against that brand's size chart — which varies wildly between a fast-fashion brand and a heritage tailoring house — is a lookup you still have to do yourself. A photorealistic render of a garment on someone "close to your build" is not a substitute for a size chart, and nothing about the render tells you it's a size down from what you'd actually need.
    • How a specific fabric actually behaves under load. A stretchy knit and a stiff woven can be rendered with equally convincing "cling" by a model that's optimized for visual plausibility, not for encoding the garment's real material properties. The image can look right for either one, even though only one of them will actually move the way the render suggests once it's on a real body.
    • Where it will actually pull or gap on your particular proportions, as opposed to the proportions of the model you happened to select. "Close to your build" is a UI choice among a finite model roster, not a match to your exact measurements — the render answers "how does this look on someone shaped roughly like me," which is a genuinely useful question, but it isn't the same question as "will this fit me."

    None of that is a knock on the technology doing something it never claimed. It's a reason to treat a try-on render the way you'd treat a well-lit product photo shot by a professional stylist: a strong signal about style and drape, and zero signal about fit tolerance.

    What a retailer owes the customer

    Two disclosure questions sit on top of this, and they're different questions.

    The first is about the image itself: an on-model photo generated by a diffusion model is synthetic media, full stop, even when the garment and the model are both real and the render is accurate. A shopper who assumes "that's a photo of that model wearing that shirt" is assuming something the retailer knows isn't quite true — it's a generated composite, however good. That's the exact case synthetic media disclosure exists to cover, and the honest version of a try-on feature marks the image as AI-generated the same way a retoning or a background swap should be marked, rather than presenting it with the same unmarked authority as a studio photograph.

    The second is about scale and consent. Under Google's merchant program, enrollment in Apparel Try-On is automatic for brands with a shopping feed and qualifying imagery — opting out requires actively contacting support, not the reverse. That's a reasonable default for a useful feature, but it also means a meaningful share of the brands whose product photos are feeding these renders may not have made an active decision to participate. A retailer that wants to do right by its customers checks both ends: confirm the feature is intentionally on, and confirm the resulting on-model images carry an AI content label a shopper would actually notice, not one buried in a terms page.

    Where a Versely workflow fits — and doesn't

    Worth being direct here: Versely does not run a garment-on-model virtual try-on model. If your workflow needs "customer picks a body type, sees the exact SKU rendered on it," that's a purpose-built capability like Google's, not something a general image or video generator approximates well — cross-attention between paired garment/body U-nets is specialized architecture, not a prompt technique.

    What Versely does carry, and where product-focused apparel brands actually use it, is the adjacent problem: generating scene and motion around a product you already photographed, without re-inventing the product itself. The AI product video generator runs on reference-to-video models, and the workflow looks like this:

    1. Shoot two or three clean reference angles of the real garment — flat-lay, on a hanger, folded — the same way you'd prep any reference-to-video job. No model, no body, just the product.
    2. Give those references to the agent along with the scene you want — a campaign backdrop, a studio you didn't rent, a season you're not shooting in. The reference-to-video model treats the garment photos as the fixed object and builds everything else around them.
    3. Check the label at 100% zoom before you ship it. Fine print is the first thing a generative model loses, and on a garment tag or a woven label, it's the one defect a buyer will actually notice.

    That's a genuinely different job from try-on: it's for the campaign shot of a product existing in a place you didn't rent a studio for, not for showing a shopper how a specific size will sit on their body. Retail and apparel teams building out a fuller content operation — lookbooks, seasonal campaigns, size-diverse lifestyle sets — will find the broader landscape of what's buildable for their category on Versely's audience hub, but the try-on decision itself still belongs to a dedicated provider, disclosed properly, with its size-chart homework done separately by the customer.