Guides

    Plastic skin is a guidance problem

    Airbrushed mannequin skin is high guidance amplifying low-frequency signal. A working CFG band, texture vocabulary, and why a fine-tune beats prompt patches.

    Versely Team8 min read

    Airbrushed mannequin skin is usually not a missing adjective. It is CFG scale turned up until the sampler amplifies the cheap, low-frequency part of the signal (broad colour, broad shading, a porcelain plane) and starves the expensive, high-frequency part (pores, vellus hair, uneven specular). The face is on-brief and dead.

    A June 2025 paper on arXiv (Song and Lai, 2506.21452) traces the oversaturated, unrealistic look of high classifier-free guidance to redundant information accumulating in low-frequency signals. That is the plastic-skin mechanism with a name. You do not need their proposed sampler to use the observation. You need to stop running portraits in the 7–15 guidance band that tutorial screenshots still copy, and you need texture words that actually survive sampling.

    This is a generation-time problem. It is adjacent to, and not the same as, the retouch failure where "smooth the skin" planes pores off an already-good face. That collapse is covered in material fidelity for skin, fabric, and metal. Here the skin never had pores to lose.

    What high guidance is doing to the face

    CFG is the weight on the difference between a prompt-conditioned prediction and an unconditioned one. Low values let the model's sense of a natural image dominate. High values exaggerate the gap. The glossary description of the high end is the right one: over-saturated colour, hard contrast, a burnt quality. On skin, the same exaggeration shows up as a wax plane. Broad shading is easy to amplify. Pore-scale noise is not, so the slider turns up the makeup and turns down the skin.

    That is also why "more photorealistic" in the prompt does not rescue a CFG-12 portrait. Guidance amplifies what the model already understood. "Photorealistic" is a category the size of the entire photographic training set. The sampler has nothing specific to push toward at the frequency that would look like skin, so it pushes saturation and smoothness instead.

    Two corollaries.

    Distilled and turbo variants want even lower guidance than their parents. They are trained to work at very low CFG. Feeding them the value that suits a full model produces the burnt look immediately. If a "fast" portrait is the most plastic of the batch, the inherited CFG is the first suspect, not the prompt.

    Flux-family models are a different instrument. They commonly run at guidance 1 with flow matching, and FLUX.2 does not expose a negative channel either. You cannot "drop CFG from 12 to 5" on an endpoint that is not doing CFG. On Flux, plastic skin is a prompting and checkpoint problem: photographic facts in the positive prompt, a photoreal model rather than an illustration one. The CFG advice in the next section is for stacks that actually expose the slider (Stable Diffusion-class checkpoints, many SDXL fine-tunes, local ComfyUI graphs).

    A working band, and how to find yours

    Copied numbers are a starting guess. The usable middle differs by model. The practitioner band that keeps showing up for full, non-distilled checkpoints is 3–7, against the 7–15 range still pasted under portrait tutorials. Below 3 the face often drifts off-brief (pretty, wrong person, wrong age). Above 7 the wax sets in.

    Find the number with a strip, not a vibe:

    1. Lock the seed and the prompt.
    2. Run the same portrait at 3, 5, 7, and 9 (and at 1–2 if the model is distilled).
    3. Zoom into a cheek or forehead at 100%. Look for pore-scale variation, not for whether the face is "sharp".
    4. Pick the lowest value that still matches the brief. Obedience above that is mostly saturation.

    If the endpoint does not expose CFG (most hosted Flux, many API image models), skip the strip. You do not have that lever. Spend the same discipline on texture vocabulary and on model choice, below.

    Do not raise CFG because the model ignored a term. The glossary pitfall is the right one: if a word meant nothing to the model, more guidance amplifies nothing. Change the word, or change the model.

    Texture vocabulary that survives sampling

    "Highly detailed skin, photorealistic, 8k, intricate pores" is a quality plea. It occupies attention and does not specify a surface. The sampler needs named, local, slightly imperfect facts at the scale of skin, because those are high-frequency instructions rather than a request for "more".

    Words that actually move a cheek:

    • visible pores along the nose and cheeks
    • faint vellus hair on the jaw
    • uneven specular highlights on the forehead
    • a dry patch beside the mouth
    • slight redness around the nostrils
    • peach fuzz at the hairline
    • film grain, candid, not studio-retouched

    Words that mostly do not:

    • photorealistic, ultra detailed, hyperreal, 8k, 4k
    • beautiful skin, flawless, baby-smooth (these actively ask for the wax)
    • cgi, airbrushed, plastic in a negative prompt (on CFG models this can help a little; on Flux it names the concept)

    A working portrait clause looks like this:

    A woman in her forties, three-quarter view, window light from camera left. Natural skin texture, visible pores, a dry patch beside the mouth, faint vellus hair on the jaw, uneven specular on the forehead. Shot on an 85mm lens, Kodak Portra 400, candid, not retouched.

    Compare the clause people actually paste:

    stunning portrait, beautiful woman, perfect skin, highly detailed, photorealistic, 8k, masterpiece

    The second prompt asks for the average of every retouched beauty still in the training set. The average of those stills is a mannequin. The first prompt gives the sampler something to do at pore scale, and it gives the colour a film stock instead of a "cinematic grade" the model will fake with saturation.

    Imperfection has to be specific. "Slightly imperfect skin" is as empty as "photorealistic". One dry patch, one region of pores, one named catch of light is enough. Three competing skin essays in the same prompt average out.

    Run this on a photoreal image model, not an illustration one. On Versely, Flux 2 Max is a realism-first option in the current catalog; the image generator is the place to A/B photoreal models on the same prompt. An illustration checkpoint will keep returning wax no matter how many pores you name, because its prior for a face is paint.

    Why a fine-tune beats another adjective

    Prompt patching fights the prior every step. A photographic fine-tune (or a LoRA trained on real portraits) is the prior. Skin statistics, including the high-frequency mess that reads as living, get baked into the weights. The prompt then has to do less heroism, which is why a mediocre prompt on a photographic checkpoint beats a careful prompt on a general illustration model.

    Two constraints, both from how adapters actually work:

    • A LoRA is bound to its base. It does not follow you to a new family. Retrain or replace it when you change checkpoints.
    • Variety in the training set is the whole game. A pack of twenty photos from one shoot teaches that day's lighting as part of "skin". Mixed angles, mixed light, mixed distances are what separate a complexion from a memorised snapshot.

    On a hosted catalog you often cannot load an arbitrary portrait LoRA. Pick the model whose prior is already photographic, and stop trying to prompt an illustration model into dermatology. If a region is the only thing failing (forehead too smooth, a cheek like glaze), a masked pass in the photo editor with the same texture clause is cheaper than a new fine-tune and cheaper than raising CFG "so it listens".

    Finish in a grade, not in the sampler. High guidance is often doing a fake grade (crushed blacks, candy midtones, glowing highlights) at the same time as it is killing skin. Generate in a flatter, more natural rendering, then grade the file. Expecting the sampler to deliver a finished look is how you get both the wax and the candy.

    FAQ

    What CFG should I start a portrait at?

    On a full, non-distilled checkpoint that exposes the slider: 5, then look at 3 and 7 with the seed locked. Stay in 3–7 unless the face is drifting off-brief at the low end. On distilled and turbo variants, start much lower, often near the model's own default. On Flux-family endpoints, there is often no slider to turn; specify texture and pick a photoreal model instead.

    I already wrote "detailed skin texture". Why is it still plastic?

    Because that phrase does not name a surface. It is a quality plea, and guidance will amplify smoothness and saturation before it amplifies pores. Replace it with two or three local facts (pores along the nose, a dry patch, uneven specular) and drop any "flawless / perfect / airbrushed" language. Then check CFG, not the next synonym for "detailed".

    Does lowering guidance make the face ignore my prompt?

    It can, which is why the band has a floor. Too low and the unconditioned prior wins: pretty, off-brief, sometimes a different person. The strip of 3 / 5 / 7 / 9 is how you find the lowest value that still obeys. Raising past that point buys saturation, not likeness.

    Can I recover pores by upscaling a plastic face?

    An upscaler will invent pore-like texture if you ask it to, and it will invent it on top of a wax plane, which reads as a skin overlay rather than as skin. Fix guidance and vocabulary at generation time. Use upscaling to enlarge a face that already has structure, not to hide a CFG problem. Two 2x passes on a decent source beat one 4x pass on a mannequin.