AI Models

    Which image models hold architectural perspective

    Bowed lines, disagreeing vanishing points and drifting facade counts sink generated architecture. The traits that help, and the prompt that fixes interiors.

    Versely Team9 min read

    Show a generated interior to someone outside the industry and they say it looks great. Show the same image to an architect and they find the problem in about four seconds, then point at the ceiling line and say it bends.

    Architecture is an unusually harsh test for image models because it is the one subject where the audience has a mental ruler. Nobody knows exactly what a face should look like. Everybody knows a wall is straight, a door is about two metres tall, and the windows on the third floor should line up with the windows on the second.

    Four failures that get a render rejected

    Bowed lines. A long straight edge develops curvature along its length. Ceiling-to-wall junctions, rooflines, the top of a kitchen island, the nosing of a stair. The bend is usually gentle and it is usually in the middle third, which is exactly where a laser level would be pointed.

    Disagreeing vanishing points. Parallel lines in the same plane should converge to a single point. Generated interiors frequently give the floor one convergence and the ceiling another, or resolve the left wall to a point the right wall never heard of. This produces the specific sensation of a room that feels wrong before you can say why, and it is the failure most likely to survive a casual review and die in a client meeting.

    Repeat-count drift on facades. A building elevation with seven windows per floor needs seven on every floor. Models produce seven, then six, then eight, with mullion counts that shift between openings and column spacing that goes irregular toward the frame edge. This is the same individuation weakness that produces duplicated faces in crowd shots, applied to modular building elements.

    Interior scale. The hardest one to fix and the easiest to miss. A ceiling that reads at four and a half metres when the brief said 2.7. A kitchen island the length of a car. Dining chairs that would seat a giant. Models have no unit system, so scale is inferred from what usually appears together, and "usually" is not a dimension.

    A fifth, less common but fatal when it appears: impossible junctions. A staircase that does not land on a floor, a beam passing through a window head, a wall meeting the ceiling at two different heights on either side of a doorway.

    Why this is structurally hard

    Diffusion models denoise toward local plausibility. Each region is resolved into something that looks correct given its neighbours, in a compressed latent space where the working representation sits well below final pixel resolution. There is no camera in that process, no scene graph, and no constraint that says "this edge is one line."

    Perspective is a global constraint. A straight edge running 900 pixels across the frame is assembled from many local decisions, none with access to the whole edge, and a two-degree deviation in the middle looks entirely plausible locally. The same logic explains the rest: counts drift because no local patch counts, and scale drifts because no local patch measures.

    This is why "make it more photorealistic" does not help. Photorealism and geometric correctness are different axes, and models are considerably further along the first.

    What predicts whether a model copes

    There is no architecture score in any public arena, and anyone publishing a definitive ranking is describing a handful of generations. What you can do is read the traits that correlate, which are visible in the catalog before you spend anything.

    Trait Why it matters here Where it shows up
    Strong prompt adherence Perspective has to be instructed, and a model that drops half your explicit constraints will drop the ones about camera height GPT Image 2 carries prompt adherence as a stated capability
    Dense-layout handling Facade repeats and mullion grids are a layout problem before they are a rendering problem Seedream 5 Pro lists dense layouts alongside typography
    An explicit reasoning pass Planning before rendering is exactly what a global constraint needs HiDream O1 and Luma Uni 1 Max both carry reasoning in their entries
    A real negative prompt field Distortion terms belong in a parameter, not in the scene description Varies by model; check the schema, not the family
    High output ceiling More pixels on small repeated elements, so mullions and balusters stay individuated Imagen 4 and GPT Image 2 both go to 4K

    Two things that do not predict it. General arena position, because Elo rewards side-by-side preference and a bent ceiling loses to nobody in a thumbnail comparison. And photographic-realism reputation, because a model like Flux 2 Max renders materials and light superbly, and that is entirely orthogonal to whether the room is square.

    One practical note on register. The prompt enhancer applies family-specific guidance by model name, and the families want different things: Flux's guidance calls for camera, lens, aperture and film-stock detail, which is exactly the vocabulary architectural correction needs. Imagen's asks for detailed natural-language coverage of subject, environment and lighting. Midjourney's wants comma-separated descriptors. Several models, including GPT Image 2, Nano Banana 2, Qwen Image and Krea 2 Medium, fall back to generic guidance, so nothing is layered on for you and you have to be explicit yourself.

    The prompt structure that stops interiors bending

    This is the payoff. Order matters, because early tokens carry more weight in practice, and the geometric constraints need to be early.

    1. Camera position and height, in real units. "Camera at 1.6 m eye height, standing in the doorway, facing the far wall."
    2. Name the perspective explicitly. "One-point perspective, far wall square to camera." This single clause is the highest-yield item on the list, because it removes the ambiguity that produces disagreeing vanishing points. If you want two-point, say two-point and say which two walls.
    3. Give a lens. "35 mm equivalent, no wide-angle distortion." Lens language constrains how much convergence is plausible, and on Flux-family models it is the register the applied guidance is already reaching for.
    4. One dimensional anchor. "2.4 m ceiling; a standard 2.1 m door on the left wall." One real measurement gives everything else something to scale against. Two is better. A room with no anchor will be sized by vibe.
    5. Count the repeats. "Five identical full-height windows along the right wall, evenly spaced." Stating the count converts a drift problem into an instruction the model can be checked against.
    6. Materials and light last. Oak floor, plaster walls, north light. These matter, and they are the part the model is already good at, so they do not need the prompt's prime real estate.
    7. Exclusions in the negative prompt field, where the model exposes one: fisheye, barrel distortion, curved walls, tilted horizon, warped ceiling. That is a separate parameter read independently, not a line of prose.
    8. Aspect ratio from the model's enum, not described in words. Writing "wide interior shot" does nothing to the frame shape; setting the value does.

    A worked example, in that order:

    Camera at 1.6 m eye height in the doorway, facing the far wall. One-point
    perspective, far wall square to camera. 35 mm equivalent, no wide-angle
    distortion. Ceiling height 2.4 m, standard 2.1 m door on the left wall.
    Five identical full-height windows evenly spaced along the right wall.
    Wide oak floorboards, white plaster walls, soft north daylight, no people.
    

    Negative prompt field: fisheye, barrel distortion, curved walls, tilted horizon, warped ceiling, duplicate windows

    Where a model exposes a batch parameter, generate several at once and compare geometry before refining further. And where a model takes style references as a separate field, put the look there and keep the written prompt on geometry, which is what the words are good at.

    Checking it in thirty seconds

    Three tests, in order, before you show anyone.

    1. The straight-edge test. Put a ruler, a guide, or the edge of a window along the longest line in the image. Ceiling junction first, then the floor line. Bowing shows immediately.
    2. The count test. Count the repeated element on every floor or every bay. If the answer differs, it is not a render, it is a sketch.
    3. The door test. Find any door and use it as a 2.1 m stick. Measure the ceiling, the island, the sofa back against it. Scale errors that were invisible become obvious the moment there is a unit in the frame.

    If a render fails one test and passes the others, inpainting the offending region is often faster than regenerating, because the rest of the image is already correct and reseeding gives you a new room. If it fails two, regenerate with the constraint stated more explicitly rather than trying to repair geometry locally.

    Studios building this into a real pipeline will find the architects page more directly relevant, and the text-to-image tool page lists what is currently in the catalog.

    FAQ

    Which single model should I standardise on for architecture?

    Do not standardise on one. Pick two with different strengths, generate the same prompt on both, and choose per project. The reason is that the failures are uncorrelated: a model that holds lines may drift counts, and a model that nails counts may size the room wrong. Versely's agent chat can fan one prompt across several named models in a single request, which makes a two-model habit cost you nothing in time.

    Do reference images fix perspective?

    They help with style and material more than geometry. A reference carries a look; it does not carry a camera. If you have an actual base photograph of the space, an image-to-image or edit path is far stronger than a reference on a fresh generation, because you are starting from geometry that is correct by construction.

    Why do exteriors work better than interiors?

    Exteriors usually have one dominant plane, more distant framing and fewer converging edges inside the frame, so there is less for the model to get wrong and less for the viewer's ruler to catch. Interiors put three planes and a floor in a single frame at close range, which is the maximum-difficulty case.

    Is upscaling worth it for architectural renders?

    For delivery, yes. As a fix, no. Upscaling adds pixels to the geometry you already have, so a bowed ceiling becomes a larger bowed ceiling. Get the geometry right at generation, then upscale for print or presentation.