Guides

    Layout-first prompting for posters and covers

    Some models now plan composition before they render. Prompt patterns that pin zones for posters, covers and thumbnails without dropping into coordinates.

    Versely Team9 min read

    There are two ways to make an image model put things where you want them. You can hand it numbers — bounding boxes, a coordinate grid, an explicit spec — or you can describe the composition in a way the model can plan against before it starts rendering pixels. Those are different techniques with different failure modes, and the second one got materially more viable this year.

    Reve 2.1, released 9 July 2026, plans layout before it renders as an explicit part of how it works. That capability changes what a prompt should look like: when a model is arranging before it draws, the highest-value thing you can put in a prompt is structure, not adjectives. Most poster and cover prompts do the opposite — they lead with subject and mood and mention placement as an afterthought, which is roughly the worst possible ordering.

    Coordinates versus planning: two different tools

    The coordinate approach is fully covered in placing type by coordinates with JSON layout prompting, and it is the right tool for a known, repeating template. This post is about the other half.

    Coordinate prompting Layout-first prompting
    You supply Numeric bounding boxes per element Zones, hierarchy and relationships in prose
    Best at Repeating a fixed template exactly One-off compositions that still need structure
    Breaks when Aspect ratio changes; boxes need redoing The model has no layout planning stage
    Failure mode Rigid, awkward, ignores good design sense Drift — plausible layout, not your layout
    Ports across models Poorly, format-specific Well, it is just clear writing

    The honest summary: coordinates give you precision at the cost of design judgment, and layout-first gives you design judgment at the cost of precision. For a poster series where every issue uses the same skeleton, coordinates win. For a cover, a one-off key art, or a thumbnail where you want the model's compositional instincts working for you inside a constraint you set, layout-first wins.

    Seven prompt patterns that pin a zone

    These are phrasings, not theory. Each one does a specific job and they compose.

    1. Declare the structure before the subject. Put the layout sentence first. Models weight early tokens heavily, and a composition described after two sentences of mood gets treated as a detail.

    Composition: three horizontal bands. Top band, upper 25% of the frame — a clear sky area reserved for a headline. Middle band, 50% — the subject. Bottom band, lower 25% — a dark, unlit surface reserved for a logo and a line of small text. Subject: a hiker in a red jacket standing on a ridge at golden hour.

    2. Give each zone a job, not a coordinate. "Reserved for a headline" tells the model what kind of emptiness you need. "Top-left" tells it a direction and nothing about the requirement.

    3. Reserve negative space as an explicit instruction. This is the single highest-leverage pattern for anything that gets type laid over it later, and it is the one people most often leave out.

    The lower-left third stays empty — flat, unlit background with no detail, no texture and no subject elements entering it.

    4. Anchor type areas to something physical in the scene. Models place objects far more reliably than they place abstract regions. A wall, a banner, a table surface or a patch of sky is a stronger anchor than a percentage.

    A plain concrete wall fills the right third of the frame, clean and evenly lit, with the subject standing to the left of it.

    5. Express hierarchy as ratios. Relative size lands better than absolute size, because the model has no canvas dimensions in mind while it plans.

    The headline area is roughly three times the height of the sub-line area beneath it.

    6. Declare the exclusion, not just the inclusion. What must not enter a region is a separate instruction from what should occupy it, and stating both narrows the outcome considerably. A negative prompt does related work on models that support one, but stated positively inside the main prompt this also survives on models that do not.

    7. State reading order. If the design has a sequence — headline, then image, then call to action — say so. It gives the planning stage something to optimise contrast and eye-path against rather than distributing emphasis evenly.

    The three-band poster template

    Combine the patterns and you get a reusable skeleton. Fill the bracketed parts:

    Composition: vertical poster, three horizontal bands. Top 25% is open sky, evenly lit, no detail, reserved for a headline. Centre 50% holds [subject], full body, centred horizontally, not crossing into the top band. Bottom 25% is a flat dark surface with no texture, reserved for a logo and one line of small type. Reading order runs top to bottom. Style: [style]. Lighting: [lighting]. Nothing enters the outer 5% margin on any edge.

    That last sentence is the one that saves reprints. A margin declaration keeps content out of the trim area, which matters the moment the file goes to print rather than to a screen — the broader version of that discipline is in AI poster and print design for business.

    For covers rather than posters, swap the band ratios: album and book covers usually want a 20/60/20 split with the type block anchored to the top, and a square aspect ratio declared explicitly rather than left to the model's default, since the same prompt at 3:2 will rearrange every band.

    Thumbnails: write the safe area into the prompt

    A YouTube thumbnail is 1280×720 and it never gets viewed clean. A duration pill sits over the bottom-right corner, a progress bar can cross the bottom edge, and in some surfaces a channel avatar or title text crowds the lower area. Designing as if the full frame is available produces a thumbnail that tests well in your editor and badly in a feed.

    The prompt fix is a single added sentence:

    Keep the bottom-right quadrant visually quiet — background only, no face, no text, no high-contrast detail. The subject's face sits in the left half of the frame.

    That is a safe-area contract expressed as composition rather than as a post-production crop, and it is far cheaper than discovering the collision after the batch is generated. The same logic generalises to every platform that overlays chrome on your image, which is the argument in designing once for 1.91:1 through 4:5. If you would rather start from a cover-shaped brief than write the constraint each time, the AI thumbnail generator gives you two routes: generate the cover as an image in its own right, composed for a small crop, or pull a still out of the finished clip and edit it into one.

    Where layout-first still drifts

    Be realistic about what this technique does not fix.

    Type rendering and type placement are separate problems. A model can plan a headline zone perfectly and still render the letters inside it badly. If the deliverable has real copy in it, the reliable route is still to generate the composition with a reserved empty zone and set the type yourself in a layout tool. That is not a workaround, it is the correct division of labour — and it is what makes pattern 3 the most valuable one on the list.

    Element count is the other limit. Two or three zones hold well. Five or six start trading against each other, and that is the point where coordinates earn their rigidity. If you are consistently past four pinned elements, go read the JSON approach.

    And prompt adherence varies sharply by model. A layout-first prompt on a model with no planning stage is just a long prompt, and it will be averaged into a plausible composition rather than yours. Test the skeleton on two or three candidates before you standardise on it: Seedream 5 Pro is a reasonable starting comparison for typography-adjacent work, and running the same skeleton through text to image across a couple of models is a ten-minute check that saves a lot of re-prompting later.

    FAQ

    Should I use prose zones or JSON coordinates?

    Use zones for one-off compositions where you want the model's design sense working inside your constraint, and coordinates for a repeating template where the same skeleton has to come out identical every time. If you are pinning more than about four elements, coordinates are the more reliable tool.

    Does putting the layout sentence first really matter?

    It matters more than most single prompt edits. Early tokens carry more weight, and a composition instruction buried after subject and style description tends to get treated as one attribute among many rather than as the frame everything else fits into. Moving it to the front is a free change worth testing on your own prompts.

    Why reserve empty space instead of asking the model to render the headline?

    Because rendering legible type inside a generated image is still the least reliable part of the job, and a mis-rendered headline means regenerating the whole image. A reserved, deliberately empty zone means the image is final and the type is a separate, editable layer you control completely.

    Will the same layout prompt work at a different aspect ratio?

    No, and this is the most common way a working prompt breaks. Band percentages that produce a good vertical poster produce a cramped or stretched composition at 16:9. Declare the aspect ratio explicitly in the prompt and re-tune the band ratios for each format rather than assuming they transfer.