Guides

    Your edit model repainted the whole image

    Instruction-based edit models regenerate the whole frame. Scope the change, then composite the region back so the background stays put.

    Versely Team9 min read

    You asked to change the mug. The mug is now a glass. The wall behind it is a different white. A picture frame has shifted a centimetre. The wood grain on the table has been invented again. Nothing in the prompt named those things, and they still moved.

    That is not the model ignoring you. Instruction editors regenerate the picture, steered by your instruction and by the input image, and they do it globally. "Leave everything else unchanged" is a hope the sampler is not built to honour as a pixel guarantee. Catalog copy will often describe these models as repainting only what you named. Treat that as the intent. The guarantee lives in a mask, or in a composite you do yourself after.

    If the background has to survive byte for byte, do not trust the instruction. Scope the change tightly, then composite the edited region back onto the original.

    Global regeneration is the default

    Edit-image models take an existing picture and a delta. They are not a layer in Photoshop. The input is encoded, noised, and decoded under the new instruction. Every pixel is eligible. Regions you did not mention tend to stay similar, which is why the tool is usable at all. Similar is not identical. On a flat wall, a small luminance shift is a new wall. On a product shot, a quietly restyled backdrop is a reshoot.

    Image-to-image is the same mechanism with an explicit strength, or denoise, dial. Low strength keeps the original nearly intact. High strength keeps a silhouette. Instruction editors hide that dial behind the prompt, which is why the global shift feels like a bug rather than a setting. Raising denoising strength because the named change is not happening is the usual way to make it worse: you give the model more freedom, it produces a different picture, and it still may not have touched the mug.

    Two ways this gets worse before the sampler even runs:

    • Re-describing the whole scene as the "edit." "A woman in a red dress standing in a park at sunset" is a new picture. "Change her dress to red, keep pose and background exactly as-is" is a delta. The first prompt hands the model a competing scene. It will reconcile by regenerating.
    • Running the prompt through AI-enhance and pasting the result. The enhance step expands a prompt into subject, setting, light, mood, palette, composition, and style. That is correct for a blank generation. On an edit, it turns "change the shirt to red" into a full new-scene description. Trim it back to a change statement before you submit.

    Flux Kontext is the one-reference case: one image in, instruction in, new picture out. Sequential single-change edits, each result fed back in as the next input, are how you keep it from inventing a second scene. GPT Image 1 Edit is built for the same job with more reference headroom. Neither one is a mask. Skip the input image and the request is rejected. Attach it and then write a full scene, and you have supplied two pictures for the model to argue about.

    Scope the instruction. That is not a mask.

    Scoping is how you reduce the area the model thinks it has permission to touch. It is still a suggestion.

    Write the delta in two clauses, always:

    1. The change, named as a single object or attribute.
    2. The hold, named as pose, background, lighting, framing, identity, product label.

    Examples that actually constrain the sampler:

    Change the mug on the table to a clear glass tumbler.
    Keep the table, wall, window, and the rest of the objects exactly as they are.
    Do not change lighting, camera, or crop.
    
    Remove the crumpled receipt in the lower-left of the table.
    Do not change the bottle, the label, or the wood grain.
    
    Replace the background with a plain light-grey sweep.
    Keep the subject, pose, edges, and the light on the subject unchanged.
    

    What does not constrain it:

    • "Make it better."
    • "Studio product shot" (that is a restyle; the whole frame is in play).
    • Three unrelated changes in one sentence. "Change the shirt, add a hat, swap the background, make her smile" is four regenerations in one pass. Chain them. One change per call.

    The photo editor and the edit-a-photo agent path both take this kind of instruction. They will look right at a glance. Zoom to 100% on a region you did not mention. If it moved, the instruction was never going to be enough.

    Flicker between the original and the edit at full opacity. Motion in the "unchanged" area is the global regenerate showing itself. If you only inspect the mug, you will ship a new room.

    Composite the region back onto the original

    This is the step that turns a global regenerate into a local edit, without waiting for a model that can actually freeze pixels.

    1. Keep the original file. Never overwrite it with the edit.
    2. Run the scoped instruction. Accept that the whole frame may have shifted.
    3. Decide the region you actually wanted: the mug, the receipt, the sky, the hand.
    4. Mask that region on the edited image, with a feathered edge and a little breathing room, the same craft as inpainting. Too tight and you get a cookie-cutter. Too loose and you bring the shifted wall back in.
    5. Composite that masked region over the original. The original is now the source of truth for everything you did not ask about.
    6. Inspect the join at 100%. If the light on the object no longer matches the original table, you either widen the mask to include the contact shadow, or you reject the edit and go to a real mask-based inpaint.

    generate_image_from_image exposes a mask parameter for models that will take one. When the model honours it, pixels outside the mask are the guarantee you wanted, and composite-back is unnecessary. When the model ignores it, or the control is not wired through for that card, composite-back is the whole method.

    Use composite-back when:

    • The named object came out right and the rest of the frame drifted.
    • A logo, barcode, legal line, or product label must not move at all.
    • You are comparing two edits of the same object and need the background identical for a fair A/B.

    Skip it when:

    • The change has to propagate: a new bottle must cast a new shadow and a new reflection. A hard composite will leave the old shadow behind, which is the failure element re-render versus a mask is built around. In that case, let the global regenerate run, then verify the regions that still must not move, and composite those back (the label, the barcode) rather than the object.
    • You are restyling the whole picture on purpose. There is no "region."

    When to stop instructing and start masking

    Instruction editing is the right default for a clearly nameable, unambiguous object, and for batch work where you will not draw forty masks. Masking is the right default the moment any of these is true:

    Situation Why instruction is not enough
    Two similar objects in frame The noun cannot disambiguate. "The red mug" with two mugs will change the wrong one, or both.
    A region that must not move, legally or commercially Only a mask, or a composite back onto the original, is a pixel guarantee.
    The background already drifted on the first try A second instruction will drift it again, differently.
    Fine edges: hair, glass, type Global regenerate plus a hard cutout is how you get a halo. Inpaint the region.
    The model keeps "helping" by restyling the set You are fighting the sampler's prior. Constrain geometry, not language.

    Practical order on a client still:

    1. One scoped instruction. Inspect at 100%, original versus edit.
    2. If only the target moved, composite it back and ship.
    3. If the target is wrong, do not add adjectives. Mask and inpaint, with a prompt that describes what belongs inside the mask, not the whole picture.
    4. If the target is right and the set must update (shadows, bounce), keep the global edit, then composite the protected regions of the original back on top.

    Pick a card that is actually in the edit-image category, not a generate model with an edit-sounding name. Then assume it will repaint the frame, and build the composite step so that does not matter.

    FAQ

    If I write "do not change anything else," why does it still change?

    Because that clause is more text in the same embedding, not a freeze. The model is not running a difference check against the input. It is generating a new image that is supposed to be close. Closeness is measured in the way the sampler was trained, which cares more about "still a kitchen" than "this exact plaster." A mask, or a composite onto the original, is how you enforce the clause.

    Does a lower denoise freeze the background?

    It reduces how far the background can walk. It also reduces how far the mug can walk, which is why the named change sometimes fails at the same setting that would have kept the wall still. Bracket low, medium, and high on a locked seed rather than guessing once. If no setting does both jobs, pick the setting that got the mug right, then paste it onto the original wall.

    Should I use inpainting from the start?

    Yes, when you already know the region, when two objects could match the noun, or when a label cannot move. Instruction editing is faster to fire and worse at guaranteeing. Fast path first, then the mask on the ones that drift. Starting on a mask for a forty-image batch is how the batch never ships; skipping the mask on a regulated pack shot is how it gets rejected.

    Why did the enhance step ruin a good edit prompt?

    It is doing the job it was built for: turning a short prompt into a full scene description. There is no edit-versus-generate switch in that rewrite. If you use it, delete the extra setting and mood language it added, and keep the single change. Pasting the expanded text is how a local edit becomes a new photograph.