Guides

    Uncrop Without Warping: Extending Backgrounds That Hold Up

    Extended edges bend straight lines, duplicate hardware and drift in colour. How much unmasked context to feed a fill model, and where the platform crops anyway.

    Versely Team8 min read

    A tightly framed product shot arrives, and the brief wants it running as a full-bleed vertical banner. Nothing about the product changes — the fix is supposed to be the space around it. That's outpainting, sometimes called uncropping, and it's a genuinely different bet than a normal edit: instead of changing something in the frame, the model is inventing territory that was never photographed at all, anchored only by whatever real pixels sit at the border. Get the anchoring wrong and the invented territory looks exactly like what it is.

    Three ways an extension breaks

    The failures aren't random — they cluster around three specific things a model has the least real information to work from. Straight lines bend because architecture and hard edges are the case with the least tolerance for error: a doorframe, a shelf edge or a horizon that's dead straight in the original photo has to stay dead straight for another few hundred pixels the model is inventing, and small angular errors that would be invisible in organic content read instantly on anything geometric. Hardware duplicates because a model extending a shot with a visible light fixture, a door handle or a repeating architectural element often solves "continue this pattern" by literally repeating the nearest instance rather than inferring what a non-repeating continuation should look like. Colour drifts because the model is estimating the extension's white balance and grade from local context, and a shift of even a few percent compounds visibly across a large extended area, especially against a smooth gradient like sky or a plain studio backdrop, where there's nothing to break up a mismatch that would hide inside texture.

    All three get worse with distance from the original pixels — the further out the extension goes, the less real information is anchoring the invention, and the more the model is working from what it expects a scene like this to contain rather than from anything actually photographed.

    How much context to feed it

    The instinct is to mask generously — cover a wide margin around the extension so the model has "room to work." That instinct is backwards for outpainting specifically. Ideogram's own fill guidance is to keep the generation window as small as possible while still covering the masked area, retaining unmasked visual reference around it rather than handing the model a wide-open canvas with only a sliver of real context to anchor against. A small window with real pixels visible at its edge gives the model something concrete to match perspective, light direction and texture against; a large window with the same amount of real context proportionally starved gives it more freedom to drift, not more room to succeed. Less unconstrained space, not more, is what produces a join that survives scrutiny.

    Resolution is the other side of the same budget. FLUX.2 edits at resolutions up to 4 megapixels while preserving detail and coherence — headroom that matters directly here, because outpainting a large area at low resolution gives the model fewer real pixels per unit of invented space to reason from. A generous resolution ceiling doesn't fix bad anchoring, but it does mean the model isn't additionally starved for detail on top of being starved for context.

    The ceiling that isn't up to the model at all

    Before extending anything, it's worth checking whether the platform you're extending for will even show the result. Pinterest truncates pins exceeding a 2:4.2 width-to-height ratio — push a pin past that, and the feed simply cuts off what doesn't fit, regardless of how clean the extension itself looks. That's a hard ceiling on how far a "make it taller" request should actually go for that specific destination: extending past a platform's own display ceiling produces a perfect result nobody sees the bottom of. Know the target ratio before the extension, not after a clean render gets silently cropped on delivery.

    Extend in stages, one axis at a time

    A single large jump asks a model to anchor a big invented area against a small strip of real border, which is exactly the setup that produces bent lines and colour drift. Extending in smaller stages gives each pass more real, already-extended pixels to anchor against than a single big jump would — each stage inherits not just the original photo but the previous stage's already-committed result as additional context. The same logic favours extending one axis at a time over pushing all four edges outward simultaneously: a model anchoring a top extension only has to reconcile with the original photo's top edge, where extending every side at once means every edge is competing for the same limited real-pixel anchor. And the two failure-prone categories from above aren't evenly distributed — skies, foliage and open water are genuinely the easy case for this technique, while architecture and anything with a straight line is the hard one, worth extra scrutiny regardless of how the rest of the extension turned out.

    A Versely walkthrough: extending a product shot to a vertical banner

    Say the source is a square product photo and the target is a 9:16 banner — extension needed on top and bottom, nothing changing in the original frame. generate_image_from_image takes the source as image_url, with mask defining the new canvas area to fill and the prompt describing only what belongs in the extension:

    "Extend this image upward and downward to fill a 9:16 frame. The product and its immediate surface must not change. Continue the studio backdrop and floor exactly as lit and coloured in the original — same white balance, same soft gradient, no new objects."

    Run it as two passes rather than one: extend upward first, generate, then extend downward from the result — each pass anchored against more real pixels than a single jump to full height would have been. Check the top and bottom thirds specifically for the three failure modes before calling it done: any straight edge in the backdrop still straight, no duplicated prop or fixture appearing near the new border, and the extended backdrop's colour matching the original rather than drifting warmer or cooler toward the edges. Versely's photo editor keeps the original alongside the result for exactly this comparison — instruction editors and outpainting extensions both re-render rather than patch a fixed region, so the honest workflow is compare-then-keep, not trust-the-first-pass.

    Checking the model before the shot

    Not every editing model handles extension work equally well, and the best AI image editing model ranking is the fastest way to see which ones on the current catalog are actually ranked for the edit-image category rather than assumed into it. Inpainting is worth knowing as the companion technique — same masked-region mechanics, opposite direction: inpainting regenerates inside a masked region while the rest of the photo holds still, where outpainting generates outside the original border into space that never existed. Confusing which one a job calls for is a fast way to mask the wrong thing.

    FAQ

    Why does a small generation window work better than a large one? A small window keeps more real, unmasked pixels visible at its edge relative to the amount of space being invented, giving the model something concrete to anchor perspective, light and texture against. A large window with the same real-context ratio gives it more room to drift, not more room to succeed.

    What's the single hardest thing to extend correctly? Architecture and straight lines — a doorframe, a shelf edge, a horizon. Small angular errors that would be invisible in an extended sky or foliage read immediately on anything with a hard, straight edge.

    Should I extend all four edges at once or one at a time? One axis at a time, generally. Extending top and bottom in separate passes — each anchored against more real, committed pixels than the last — produces a more reliable join than asking one pass to reconcile all four edges simultaneously against a single, smaller strip of original photo.

    Does the platform I'm posting to affect how far I should extend? Yes, directly. Pinterest's 2:4.2 truncation ceiling is a concrete example: extending past what a destination platform will actually display produces a clean result with an invisible bottom half. Check the target ratio before generating, not after.

    Outpainting is a bet against the edges of a photo, not against the product in the middle of it — and the tell for whether the bet paid off is always at the border: the straight line, the fixture that shouldn't have a twin, the gradient that shifted half a shade. Check there first.