Blur garbled signage into depth of field
Background text a model got wrong is not repairable, but it is not necessary either. A soft-focus pass that reads as optics, and where it stops working.
The shot is right. Lighting is right, the subject holds, the composition works. Then somebody looks at the shopfront eleven metres behind the subject's left shoulder and reads "RESTAURANI", and the whole frame collapses into an obvious generation.
The instinct is to fix the word. That is the wrong instinct, and understanding why saves a lot of wasted attempts: the sign is not repairable at that size, and it was never supposed to be readable in the first place.
Why the sign cannot be repaired in place
Text generation fails differently from every other artifact. Models learn the shapes of glyphs from images, not the letterforms of an alphabet as a system with rules, so what comes out is something that resembles writing at a glance and dissolves under attention. That behaviour is well documented across model families and is the same reason a book spine in the background says something almost-but-not-quite.
Two properties make incidental signage worse than most defects:
It occupies almost no pixels. A background sign might be forty pixels tall. That is the same small-region problem that makes hands fail — the model is solving a tiny fraction of the frame and there is very little budget spent getting it right. Regenerate that region and you will get a different wrong word.
Correctness is discrete. A hand can be plausible. A word cannot. It is either the right string of characters or it is a mistake, and there is no partial credit. So even a good inpaint has to land exactly, at forty pixels, in a region where the model has no strong prior about what the sign should say.
The whole-image reroll is worse, because you are trading a composition that works for a lottery ticket on one background element.
Blur is optics, not a cover-up
Here is the part that makes this a legitimate finish rather than a hack. A real lens renders exactly one plane sharp. Everything in front of and behind that plane falls off, and how fast it falls off depends on aperture and focal length. If your subject was photographed at portrait focal length with the aperture open, a sign eleven metres further back would be an unreadable soft shape in the real photograph. Not blurred by an editor. Blurred by physics.
Which means the model rendering it crisply is itself an error. Generated frames are frequently sharp across depths that no single lens could hold at once, and the garbled text is just the most visible symptom of that. Softening the background is not hiding a mistake. It is correcting a depth inconsistency that was already there, and the text problem disappears as a side effect.
That reframing matters because it tells you how to do the pass. You are not applying blur to a sign. You are restoring a depth-of-field falloff to the whole frame, and the sign happens to live in a band that should have been soft.
The pass, step by step
- Decide where the focal plane sits. Look at what else in the frame is genuinely sharp — individual hair strands, fabric weave, an eyelash — and take that as the subject distance. Everything gets graded relative to it.
- Mask by depth band, not by object. This is the single most common mistake. Blurring just the sign leaves a soft rectangle floating inside a sharp wall, which reads as worse than the garbled text did. Select the whole plane the sign belongs to: the building, the street behind it, the parked cars at that distance.
- Feather far more than feels necessary. The feather is the entire trick. A hard-edged mask announces itself instantly. Real falloff is gradual across depth, so the transition zone should be wide — often a good fraction of the distance between planes.
- Graduate it. Two or three bands minimum, each softer than the last. One flat blur value across the whole background is the look of a phone portrait mode getting it wrong, which is a tell in its own right.
- Use a lens blur rather than a Gaussian where there are highlights. Real defocus turns point highlights into discs with defined edges. A Gaussian smears them into grey mush. If there are practicals, streetlights or specular hits in the region you are softening, the difference is visible.
- Put the grain back. Blur removes noise along with detail, so a blurred patch inside a grainy frame reads as a smudge even when the geometry is perfect. Match the grain of the sharp areas into the softened region.
- Check the subject's silhouette for halos. A faint bright fringe tracing the subject's edge, where the background got blurred across the boundary, is the classic giveaway. Pull the mask in slightly at the edge rather than out.
For images, a masked edit can do some of this work directly: mask the background band and prompt the region as defocused rather than as legible signage. Results vary by model, and denoising strength is the dial that decides whether you get a softened version of what was there or entirely new background content. Check that you did not gain a new building. The AI photo editor is the direct surface for this; editing a photo through the agent is the route if you would rather describe the change than choose the model and the strength yourself.
Prevent it at generation time instead
The cheapest version of this pass is not doing it. Depth-of-field language is not decorative in current models — shallow and deep depth of field, along with focal length, genuinely shift what the render does, and stating them costs nothing.
Medium shot, 85mm, shallow depth of field, focus on the subject's eyes.
The street behind her falls completely out of focus; background storefronts
read as soft coloured shapes with no legible detail.
Two things are doing work there. "Focus on the subject's eyes" names the plane. "No legible detail" describes the desired state of the background rather than instructing the model not to render text, which is the phrasing that actually holds up. Telling a model "no text" tends to put text in the frame, because the noun is still in the prompt.
Framing helps too. A tighter shot, a lower angle that puts sky behind the subject instead of shopfronts, or a background wall rather than a street all remove the opportunity before it exists. Hands, teeth and signage covers the shot-design version of this for all three of the surfaces that still expose generated footage.
Where the trick stops working
The sign shares the subject's plane. If it is at the same distance as the thing in focus, blurring it contradicts the optics of the rest of the frame and every viewer feels it, whether or not they can say why. This one has no post fix. Reframe or regenerate.
The sign is large in frame. Beyond roughly a fifth of the frame width, softening stops reading as depth and starts reading as censorship. The eye has enough of the shape to know something was removed.
The text is load-bearing. A product label, a brand name, a headline — anything the viewer is supposed to read is not a blur candidate at all. That is an overlay job: generate the scene without it and add real, controlled type afterwards. Adding a text overlay is the video route, and signage and trade-show graphics from generated images covers the larger case where the graphics are the deliverable.
The frame has signs at many depths. If softening the bad ones means softening most of the image, you are not finishing a shot, you are admitting the shot is wrong. Regenerate with better framing.
Anything a reviewer has to sign off. If a brand or legal check tests whether text appears at all rather than whether anyone can read it, softening does not clear the check. Find out which test applies before you spend the pass.
Video, unless you track it. A static mask over a moving frame slides off within a second, and the blur boundary starts crawling across objects it does not belong to. The region has to be isolated properly first, which means video segmentation rather than a drawn shape — Sam 3 Video Segment produces that tracked mask, and the softening then happens against it in your editor. That is enough setup to argue for catching signage at the prompt stage on anything that moves.
FAQ
Will an upscale fix garbled background text?
No, and it usually makes it worse. Upscalers sharpen and reconstruct what is present; they have no way to know the intended string, so a confident upscaler will render the wrong word more crisply than the original did. If you are planning both passes, soften the region before you upscale, not after.
Does blurring hurt perceived quality?
It usually improves it. Uniform edge-to-edge sharpness is one of the reliable signatures of generated imagery, because real optics do not behave that way. A frame with a clear focal plane and genuine falloff reads as more photographic, not less detailed. The exception is a shot that was meant to be deep-focus, where softening a band creates an inconsistency of its own.
Can I just crop the sign out?
Often, and it is the fastest fix when the sign sits near a frame edge. The cost is composition: cropping changes your framing and your aspect ratio, and on vertical deliverables there is rarely spare margin. Worth checking first anyway, because it takes ten seconds to find out.
How much blur is right?
Enough that no character shape is recoverable, and no more. If you can still make out that something is text, the eye keeps trying to read it and the artifact survives. If the background has become a featureless wash, you have gone past what any lens does and lost the depth cue you were trying to create. Soft coloured shapes with recognisable structure but no legible detail is the target.