Comparisons

    Prompt adherence vs aesthetic defaults

    Complaints that a model ignored the prompt are usually a casting error. Two prompts at three seeds each sort image models into the two camps that matter.

    Versely Team8 min read

    Two image models can be equally good at making a nice picture and still disagree completely about whether your prompt was a specification or a suggestion. That disagreement is the single most useful thing to know about a model before you cast it on a job, and it is almost never what the model card leads with.

    Two different failures wear the same complaint

    "The model ignored me" gets said about two opposite problems.

    The first is a model that did exactly what you asked and the result is plain. Three mugs on a shelf, correctly counted, correctly coloured, lit like a stock photo from 2014. Nothing was ignored. The model simply has no opinion of its own to contribute, and you were relying on it having one.

    The second is a model that returned something better-looking than you asked for and quietly deleted half your brief on the way. You wrote "flat, even lighting, plain grey background" and got a rim-lit portrait with a shallow lens and a warm grade. It didn't fail. It overrode you, because its aesthetic default outranked your instruction.

    Versely's glossary puts the axis precisely: prompt adherence is how faithfully a model does what the prompt actually said, as opposed to producing something attractive in the same neighbourhood. It also names where the two camps diverge most visibly — counting, spatial relations, text rendering, negation, and binding several attributes to several subjects. Those are the checks that separate a model with an opinion from a model with a specification.

    Model descriptions track the split more honestly than you'd expect if you read them as engineering notes rather than marketing. GPT Image 2's catalog entry leads with strong prompt adherence. Seedream 5.0 Pro's leads with deep-thinking prompt understanding, native text rendering across 14 languages, and precise control over dense layouts. Midjourney V7's leads with photorealistic and artistic outputs. Recraft 4.1's leads with strong style control and consistency. Those are two different things being optimised, described in two different vocabularies.

    The classification test: two prompts, three seeds

    You can sort any image model in about ten minutes with two prompts. Both are deliberately unpleasant to look at, which is the point.

    Prompt A — the specification. No style words at all. Every constraint checkable by counting or reading.

    Three ceramic mugs in a row on a white shelf. The left mug is matte red.
    The middle mug is white with the word BLOOM printed on it in black.
    The right mug is tipped on its side and empty.
    

    Five binary checks: exactly three mugs, left one red, middle one white, BLOOM legible and correctly spelled, right one tipped. Count them. Do not look at the picture and form an impression.

    Prompt B — the override. Deliberately unglamorous, and it asks for the absence of everything a house look wants to add.

    A flat, evenly lit ID photo of a woman facing the camera against a plain
    grey wall. Even front lighting, everything in focus, neutral colour,
    centred, no styling.
    

    Score what came back anyway: rim light, background falloff, a shallow lens look, a warm or teal grade, styled hair, a tilt. Each one is the model spending your prompt budget on its own taste.

    Run each prompt at three different seeds. This is the step people skip and it carries most of the signal. A seed decides the starting noise, so a single frame cannot tell you whether the rim light came from the model or from a lucky roll. What survives a seed change is the model's default. Three ID photos that all arrived with the same rim light is a house look. Three that all arrived flat is an instruction-follower.

    One setup note: turn off any prompt expansion or enhance step for this test. An expander rewrites your input before the generator ever sees it, which means you stop being the author of the thing you are grading, and two runs of the same short prompt can be expanded differently for reasons the seed cannot explain.

    Reading the scores

    Prompt A hits Prompt B override Class Cast it for
    4–5 of 5 little or none Instruction-follower Packshots, layouts, on-image text, counts, exact colours
    4–5 of 5 heavy Controllable stylist Rare. Make it your default and stop testing
    0–2 of 5 heavy House-look model Key visuals, thumbnails, editorial, anything briefed as a feeling
    0–2 of 5 little or none Neither Not a casting problem. Retire it

    A house-look model is not the worse option. It is carrying a decision you would otherwise have to make yourself in adjectives, and it makes that decision more consistently across a set than you will. Versely's glossary makes the same argument about named style presets: six images that each describe their own look in prose drift apart, and six on one preset share a grade, which is most of what makes a set read as a set. A house look is a preset you did not get to select. That is a cost on spec work and a benefit on mood work, and which one it is depends entirely on the brief.

    The per-job rule

    Before casting, answer one question about the brief: is it a specification or a mood?

    Specification means any constraint a stakeholder can verify by pointing at the image. How many. Which side. What the label says. Which hand holds the product. The exact brand colour. Cast an instruction-follower, accept that the first pass may look flat, and add style in a second step.

    Mood means the brief is "premium", "warm", "chaotic energy", "expensive but not shiny". Cast a house-look model and let the house look do the styling work you were about to type badly.

    Mixed briefs are the trap, and most real briefs are mixed. Split them rather than asking one prompt to win both fights: generate for the specification first, then restyle the winner. The reflex to avoid is escalation. Adding adjectives to a losing prompt on a house-look model makes it more confident, not more obedient, and the same dynamic shows up in camera instructions video models silently ignore.

    Two procedural rules that stop the test from lying to you:

    1. Never change model and prompt in the same step. If you switch casting, hold the prompt fixed for one run so you are grading the model rather than your own rewrite.
    2. Keep two winners, not a favourite. One instruction-follower and one house-look model, re-tested whenever either ships a new version, because a version bump can move a model across the line.

    The fastest way to run the whole thing is to fan one prompt across several named models in a single request, which the agent chat does natively. Two messages, six candidates, thirty-six frames, and a bench that holds for the next quarter of briefs.

    FAQ

    Does a more expensive model follow prompts better?

    No, and treating credit cost as a proxy for obedience is how people end up paying more for the same override. Credit cost tracks what the model renders — resolution, output size, whether extra passes run — not how literally it reads a sentence. The quality-per-credit frontier report is a better place to look for value, and the two-prompt test above is the only place to look for adherence.

    Can I force adherence with words like "exactly" or "must"?

    Emphasis is not a control channel. Models weight the content of a clause, not its urgency, and stacking intensifiers usually just makes the prompt longer. Where a model exposes a real schema field, use the field: a dedicated negative-prompt input for exclusions, a style-reference input for a look you would otherwise describe in adjectives. A parameter is a hard input. An adverb is a hope.

    Should I leave prompt enhancement on for real work?

    Leave it on for mood work, where filling in lens, light and texture usually improves a thin prompt. Turn it off for specification work, where the enhancer can invent and insert a detail you deliberately omitted, and where you need the prompt you graded to be the prompt that ran.

    Does this test work on video models?

    The adherence half does. Counting, spatial relations and attribute binding fail the same way in a first frame as in a still. The override half needs a second axis, because video models have motion defaults as well as look defaults, and a model that adds an unrequested slow push-in is doing the same thing as one that adds unrequested rim light. Score both, and score them across seeds for the same reason.