Background faces melt in wide shots
Faces occupying few pixels get invented, not drawn. Prompt shallow depth of field, crop-upscale-and-composite, or reframe so the crowd never resolves.
A face that occupies twenty pixels is not a small portrait. It is a region the model never had enough budget to draw, so it invents a face-shaped smear and moves on. Pause a wide shot and you see it: the hero is sharp, the street is plausible, and every extra in the mid-ground has a mouth that will not commit to being a mouth.
This is the same pixel-share problem that wrecks hands, pointed at a different surface. Training objectives average error across the image. A crowd face is a rounding error in that average. The model can be confidently wrong there and still score well, which is why re-rolling the wide keeps giving you a new hero and a new set of melted extras.
Do not ask an upscaler to rescue those faces. An upscaler invents plausible high-frequency detail; on a distant face that detail is a reconstructed identity, not a recovered one. Upscaling is excellent at texture and unreliable at information. A crowd face is information you never encoded.
Three fixes work. Prompt the background out of focus so the model is asked for bokeh rather than anatomy. Crop the one extra you actually need, generate that region at portrait scale, and composite it back. Or reframe so the crowd never has to resolve.
Why a crowd face is not a tiny portrait
Resolution is not a slider that adds the same quality everywhere. Pixel count is a budget. In a close-up that budget is spent on one set of features. In a wide it is spread across a street, a sky, signage, and twenty heads. Each head gets a handful of pixels. The model then samples something face-shaped from a distribution of faces, rather than rendering a particular person.
That is why the failure looks like melting rather than blur. Blur would keep a structure and lose high frequencies. Melting is structure that was never decided. Eyes drift toward each other. A jaw becomes a gradient. Two extras share a cheek. If you zoom in and the "face" looks like a different person at every zoom level, you are looking at invention, not compression.
The same mechanism shows up on other small, high-scrutiny surfaces. Hands, teeth and signage fail because they occupy little of the frame and because viewers notice them. Background faces sit in that family. They are just easier to miss in a scrub, which is why they survive review and die in the pause.
Two consequences follow. A better prompt about "photoreal people in the background" does not allocate more pixels to those people. And generating the same wide at a higher output size often fails the same way, or worse: you get a larger smear. Size without a change in composition does not convert a twenty-pixel face into a portrait.
Three fixes, in the order you should try them
| Fix | Use when | What it does |
|---|---|---|
| Prompt shallow depth of field | The crowd is atmosphere, not cast | The model is asked for unresolvable faces |
| Crop, generate large, composite | One background person must be a specific face | That face becomes most of a frame, then returns to the wide |
| Reframe so the crowd never resolves | You do not need readable extras | Silhouettes, a tighter crop, motion, or a wall instead of a street |
Start at the top. Most wides do not need a cast of extras. They need a sense that a place is occupied. Shallow depth of field, a longer lens in the prompt, and a background of "figures, out of focus" will beat "a crowded street of distinct people," because you stopped requesting anatomy the model cannot fund.
A working prompt delta:
85mm, f/1.8, subject sharp, background figures
as soft bokeh discs, no readable faces behind
the subject, shallow depth of field
Write the optical result, not the negative. "No faces in the background" names faces, and negation is a weak channel. "Soft bokeh, unreadable figures" describes the state you want. Put the lens language at the end of the prompt.
If the extra is actually a character — a colleague at the next desk, someone the story has named — skip the bokeh trick. You cannot fake a specific person with discs of light. Go to the crop.
If neither applies, reframe. A medium shot against a wall, extras in silhouette against a window, or a pan that never lets the mid-ground settle: those are cheaper than a composite. If the story does not need those people to resolve, do not spend the model's budget pretending they will.
The crop-upscale-composite loop
This is the loop for the one extra you cannot lose. It is not a finish pass on the whole wide.
- Generate the wide for composition, not for extras. Lock the hero, the set, the light. Ignore melted faces on this pass.
- Crop the extra plus context. Include shoulders and a slice of what they stand against. A face-only crop composites like a sticker.
- Generate or edit that crop at a scale where the face is most of the frame. You have converted a pixel-share failure into a portrait problem, which models can solve. Use the photo editor with a mask over the face and a short prompt: age, expression, gaze, light direction matching the wide.
- Keep denoising strength in the middle. Too low and the melt survives under a new texture. Too high and you get a different person who no longer matches the body in the wide.
- Composite with a feathered edge, matching grain and colour. If the extra looks shot on a different camera, drop contrast slightly and re-check the light direction.
- Do not upscale the wide to "sharpen the crowd." Whole-frame upscaling will invent reconstructed faces on extras and letter-shaped marks on signs. If the hero is soft, tighten the framing or upscale the hero region. Leave the crowd as bokeh or as the one composited face.
Inpainting without an explicit crop is the same principle if your editor already zooms the masked region: the extra has to occupy most of the model's budget for one pass, then return to being small. In a crowd, naming "the person on the left" is ambiguous. Mask.
A prompt that works on the cropped extra, assuming window light from camera left:
A woman, mid-30s, looking down at a phone held
in both hands, three-quarter view, window light
from camera left, same warm interior, sharp
eyes, closed mouth, no smile
Age band, action, angle, light direction, expression. "A realistic extra" gives the model the same under-determined face it invented in the wide.
If several extras must resolve, you are no longer in a wide-shot workflow. Generate the group as a medium, or generate each readable person as a plate and composite. There is no honest version of "a sharp crowd of twenty distinct faces" at a single generation's budget.
When to stop asking the model for a crowd
A 9:16 establishing wide of a market, every stall in focus, every shopper's face readable, is a request for dozens of portraits the model will not fund. Change the shot.
- Longer lens, tighter crop. The street is implied by a slice of it. One extra, maybe two, and bokeh where the rest of the crowd would have been.
- Backs and silhouettes. People walking away, or standing against a bright window. A silhouette does not have to be a face.
- Motion. A short pan, extras crossing. Moving crowds hide the melt that a locked-off wide puts on trial.
- Fewer bodies, more set. Three people in a cafe read as a cafe. Thirty melted faces read as a generation.
If you are generating stills as first frames for video, this matters twice. Image-to-video will animate whatever structure is in frame zero. A melted extra becomes a melting extra over time. Clean the still, or keep those faces out of focus, before you turn the photo into a clip. Video models do not spend their budget on background anatomy either.
Review wides paused, at 100 percent, on the mid-ground. Melted extras disappear in a 320-pixel preview and announce themselves on a phone held still.
FAQ
Will generating the wide at 4K fix the melted extras?
Usually not. A higher output size spreads the same composition over more pixels. If a face was twenty pixels of structure at 1080p, four times the pixels can still be four times a smear, or a reconstructed face an upscaler invented. Change how much of the frame that face occupies, or stop requiring it to resolve. Generate at native size and upscale is the right finishing move for texture, not for crowd anatomy.
Can I inpaint every extra in a busy street?
You can. You should not. Each extra is a portrait job plus a composite, and the light has to match on every one. Past two or three readable faces, reframe. Atmosphere does not need a cast list.
Why does shallow depth of field still leave one sharp, melted face?
The model kept someone in the plane of focus. Name the plane: "only the foreground subject is sharp." If a specific extra keeps landing sharp, mask that extra and paint them as bokeh, or crop them out of the composition. Depth of field in a prompt is a request, not a lens.
Should I upscale the wide and then inpaint?
Inpaint first, or crop-and-replace first. Upscaling bakes invented faces into the file, and you will then be inpainting reconstructions rather than smears. Fix structure at the scale you generated, then upscale the finished still if the delivery needs the extra pixels.