Guides

    Ghost Mannequin Apparel Shots Without the Mannequin

    The ghost mannequin look isn't about erasing a form, it's about what the eye reads as a hollow, worn garment — collar depth, shoulder weight, interior shadow.

    Versely Team7 min read

    Ghost mannequin photography — a garment shot so it looks worn by an invisible body, hollow but shaped, with no model or mannequin in frame — has always been a physical production trick before it was a photography style. A studio shoots the garment on a real mannequin or dress form, then shoots the interior neckline and cuffs separately with the form partly removed, and a retoucher composites the two so the collar shows depth and shadow where a head would have been but no mannequin is visible anywhere. Two shots, one composite, done by hand.

    Doing this with a generative edit model skips the physical rig but not the judgment. The output still has to convince a shopper's eye that a body used to be there, and that read comes from a short, specific list of cues — not from "remove the mannequin" as a single instruction.

    What the eye is actually checking

    Nobody looking at a ghost mannequin product photo consciously inventories collar shadow and shoulder slope. But strip any one of these out and the image reads as fake within a glance, because the eye is checking all of them at once:

    • Interior collar depth. A worn neckline isn't a flat opening — you can see a sliver of the garment's inside, in shadow, curving away from the camera. A neckline generated as a flat cutout, evenly lit all the way through, is the single fastest tell that nothing was ever inside the shirt.
    • Shoulder weight. Worn shoulders slope and settle with a slight downward curve from the collar. A garment photographed or generated flat, then just masked into a torso shape, keeps its flat-lay shoulder line — square, unweighted — and that mismatch reads as wrong even to someone who couldn't say why.
    • Cuff and hem hollow. Sleeves and hems need the same interior-shadow treatment as the collar: a visible, shadowed opening, not a solid cap.
    • A floating shadow, not a resting one. Ghost mannequin shots use a soft, close contact shadow directly under the garment, not the longer cast shadow of an object sitting on a surface — the garment should read as suspended in the exact shape a torso would hold it.

    This is the same underlying problem Google's shopping team describes solving for its apparel try-on feature, a different product from ghost mannequin work but one aimed at the identical physical target: rendering "how something drapes, folds, clings, stretches and wrinkles" convincingly on a body that isn't literally there in the source photo. Google trained that model on paired images of the same garment across two different poses, so it learns how one drape maps onto another rather than painting fabric texture from scratch. Versely doesn't carry that specific Google model — it's a shopping-side try-on feature, not a listed catalog entry — but the target it's solving for is exactly the list above, and it's the right way to think about what an edit-image model needs to be told, not just "remove the mannequin."

    The sequence, not just the instruction

    A single prompt like "remove the mannequin" asks a model to both erase a thing and correctly reconstruct four separate physical cues in its absence, which is a lot to load onto one instruction. Breaking it into an ordered pass, using inpainting at each stage, gets a cleaner result:

    1. Start from a garment already on a form — a real mannequin, dress form, or flat-lay shot with the garment's actual shape visible. Generating a ghost mannequin look from nothing but a flat-lay photo with no dimensional information at all is the hardest version of this task; give the model a real starting shape whenever you can.
    2. Mask and inpaint the visible mannequin or form, prompting specifically for what should replace it — a soft contact shadow and the garment's own interior fabric — rather than just "empty" or "gone." An empty mask with no replacement instruction tends to flatten the collar along with the form.
    3. Mask the neckline and cuffs separately and prompt for interior depth: a shadowed fold, fabric curving inward, the inside of the collar visible. This is a smaller, more targeted mask than step 2, and it's the step that actually produces the "worn" read.
    4. Check the shoulder line last. If it's still flat after the first two passes, a third targeted inpaint on just the shoulder seam — prompting for a slight downward slope and natural settle — fixes it without touching anything already working elsewhere in the frame.

    Masking generously around each area rather than tracing it tightly gives the model somewhere to blend the new shadow into the existing fabric, which is the difference between a seam that survives a shopper zooming in and one that doesn't.

    Picking a model for the pass

    Versely's catalog carries 42 models in the edit-image category, spanning instruction-based editors from Google, ByteDance, Qwen, Flux and more — the full current roster is on /models. Four are a reasonable starting shelf for a ghost mannequin pass specifically:

    • Nano Banana Pro Edit — a professional-tier editor with headroom for the kind of multi-region masking this technique needs across collar, cuffs and shoulder in one session.
    • Qwen Image Edit 2511 — the cheapest of the four per edit, and a fast way to iterate through several mask-and-prompt passes before committing to a final version on a pricier model.
    • Flux Kontext — built around context-aware generation, useful for keeping the rest of the garment's texture and color consistent while a masked region is being reworked.
    • Wan 2.7 Pro Edit — accepts up to 4 reference images per edit, worth keeping in the kit if you have real ghost mannequin photography from a past shoot you want the model to match stylistically, not just a text description of the look.

    None of these need to be "the" model for the whole job — a common pattern is cheap-and-fast for the first two or three mask passes while you find the right prompt wording, then a pass on a higher-tier editor for the final image that actually ships.

    Walkthrough: a ghost mannequin pass in Versely

    1. Open the source photo — garment on a real form or mannequin — in Versely's AI photo editor.
    2. Mask the visible form at the neckline and prompt for what should be there instead: interior fabric shadow and a soft contact shadow beneath the garment, not just removal.
    3. Generate, then mask the neckline more tightly and prompt specifically for interior collar depth — a shadowed fold curving inward — if the first pass left it flat.
    4. Repeat the same targeted mask-and-prompt approach on the cuffs and hem.
    5. Step back and check the shoulder line. If it's still reading flat, one more masked pass with a "slight natural slope" instruction is usually enough.
    6. Compare the result against a real ghost mannequin reference if you have one — the floating-shadow test is the fastest final check: does the garment look suspended in a body's shape, or does it look like it's resting on something?

    Why the sequence beats the one-shot prompt

    Every one of the four cues above is small enough to miss individually and obvious enough to notice together, which is exactly why "remove the mannequin" as a single instruction underperforms a short ordered pass. Each targeted mask only has to solve one physical detail at a time, and the model has an easier job matching light and shadow to a small region than reconstructing an entire garment's worn shape from one broad edit. The four-step sequence costs a few more generations than a single prompt — it's still faster than the two-shot physical composite it's replacing, and it's the version that actually survives a shopper zooming in on the collar.