Comparisons

    Short prompts or long prompts: what models want

    A 12-word prompt and a 120-word one fail differently. Per-class length targets, plus the ordering rule that keeps your one non-negotiable clause alive.

    Versely Team9 min read

    Write twelve words and half the models in the catalog will invent the other hundred for you. Write a hundred and twenty and a different half will quietly drop a third of them. There is no universal answer to prompt length, but there is a per-class answer, and it is more specific than "be detailed" or "keep it simple."

    The test that makes the difference visible

    Take one subject and write it twice.

    The 12-word version:

    A woman drinking coffee at a kitchen window in morning light.
    

    The 120-word version: the same scene with subject, action, environment, lighting, colour palette, composition, mood and medium all stated, plus three hard constraints buried in the middle — she holds the mug in her left hand, there are exactly two plants on the sill, and the mug is matte black.

    Run both, lock the seed so wording is the only variable, and then score them the same way: clauses landed divided by clauses written. Not "which looks better." The 120-word version will almost always look better, because it fed the model more to work with. The question is whether clause 9 arrived, and whether adding clauses 9 through 15 cost you clauses 1 through 8.

    That ratio is the useful number, and it moves in a direction most people find surprising. Past a class's ceiling, the extra clause doesn't just fail to land, it degrades the hit rate on the clauses that were landing before it. Every additional clause competes with every other clause for the same fixed processing budget, which is the same mechanic behind prompt adherence collapsing near a supplied keyframe.

    Length targets by class

    Versely's own AI-enhance step is a useful calibration point here, because it encodes a target rather than a preference: it rewrites an image prompt toward 50–150 words and explicitly instructs the rewriter to cover subject, environment, lighting, mood, colour palette, composition and style. Seven slots, 50–150 words. That is roughly 7 to 20 words per slot, and it is a reasonable band to hit by hand even if you never touch the enhancer.

    The band shifts by family, and the model-specific tips the enhancer injects tell you which direction:

    Class Target shape Why
    Text-to-image, general 50–150 words, all seven slots The enhancer's own target; covers the checklist without padding
    Midjourney V7 / Niji 6 Shorter and denser, comma-separated descriptors Its tip calls for artistic references, lighting and medium descriptors, and explicitly says to avoid generic filler words
    Flux family, photoreal Top of the band, natural language Its tip asks for camera details: lens type, aperture, film stock
    Imagen 4 Top of the band, natural language Its tip asks for detailed prompts, explicit about what you want in the scene
    Wan text-to-video One or two sentences plus a style keyword The shortest tip in the family list: straightforward scene descriptions and style keywords, and no camera or style field in the schema to write toward
    Kling / VEO / Sora Dense and camera-forward All three tips converge on naming the move: dolly, crane, tracking, pan, and temporal flow
    Seedance / MiniMax Medium, motion or narrative led Neither tip asks for camera hardware. Lead with what moves and how the scene unfolds
    Instruction image edit Shortest of everything: one change The input image is ground truth; the prompt is a delta against it

    The edit row is the one that trips people who learned prompting on text-to-image. On an instruction-edit model, re-describing the whole target picture is the single biggest failure mode. "A woman in a red dress standing in a park at sunset" hands the model a second, competing scene to reconcile against the reference. "Change her dress to red, keep everything else the same" hands it a delta. That is a 9-word prompt beating a 40-word one on the same job, decisively.

    It is also where the enhancer needs supervision. Versely's enhance step branches on image versus video, not on edit versus generate, so running "change the shirt to red" through it can hand back a full new-scene description. If you enhance an edit prompt, trim the result back to a change statement before you submit it. The per-model prompting guides note this per family rather than as a general rule, because the correct behaviour genuinely differs by class.

    The ordering rule

    Length is the easier half of the problem. Position is the half that decides whether your one non-negotiable clause survives.

    Long-context language processing has a documented weak spot in the middle of a sequence, often discussed as the "lost in the middle" effect: material at the start and the end of a long input gets attended to more reliably than material buried between them. Prompts are inputs. A hard constraint written as clause 7 of 14 is sitting in the worst position available.

    So the rule is mechanical:

    1. Open with the clause that must land. If the mug has to be in her left hand, that is the first thing in the prompt, before subject description, before setting, before anything atmospheric.
    2. Put atmosphere in the middle. Mood, palette and grade are the clauses you can afford to lose, because losing them produces something different rather than something wrong.
    3. Close by restating the constraint, once, in different words. Not as emphasis — as a second placement in a position that gets read. "Left hand" at the top, "the mug is in her left hand, label facing camera" at the bottom.
    4. Never stack two hard constraints in the same clause. Split them so each gets its own position, and if you have more than three, you have a job for reference images or a schema field rather than for prose.

    That last point matters more than the ordering. Prose is the weakest way to state a constraint a model exposes a real input for. A dedicated negative-prompt field beats "no text, no watermark" written into the positive prompt, because a negation in the positive prompt still puts the concept into the mix and models render it anyway. A style-reference image beats four sentences of style adjectives. Length can only ever compensate for the absence of a better channel.

    Where prompt expansion sits in this

    If you write 12 words on a model with provider-side expansion enabled, you are not running a 12-word prompt. You are running whatever the expander wrote, which is typically sixty words, filled in with lens, light, texture and mood it inferred rather than you specified. For casual work that is usually an upgrade, because most people write far too little. For anything you intend to reproduce, it is a problem: two runs of the same short prompt can be expanded differently, so results vary in ways the seed cannot explain, and a detail you deliberately left out may be invented and inserted.

    The practical split: short prompts plus expansion for exploration, where variance is a feature. Explicit prompts at the class's target length with expansion off for production, where variance is a defect. And if you are comparing two models, expansion has to be off on both, or you are grading two expanders rather than two generators — the same discipline that makes running one prompt across several models worth anything.

    A working procedure

    1. Identify the class first, not the model. Text-to-image, text-to-video, instruction edit and reference-driven generation have different targets, and a model can sit in more than one class depending on whether you attached a reference image.
    2. Write to the class target, using the table above. For general text-to-image, the seven-slot checklist is a fast way to hit 50–150 words without padding.
    3. Put the non-negotiable first and last. Everything else can float.
    4. Score by clauses landed, not by preference. Write the count down. Six of nine on a 120-word prompt is worse than five of six on a 60-word one, and it will not feel that way when you are looking at the pictures.
    5. Cut, don't add, when a clause fails. The reflex to lengthen is the wrong one. Removing the three softest clauses is more likely to rescue the hard one than adding a fourth restatement of it. Then push the survivor through text-to-image or the video tool at full settings once it is behaving.

    FAQ

    Is there a hard token limit I'm hitting?

    Sometimes, but the failure you are usually seeing is not truncation. Truncation is abrupt: the tail of the prompt is simply gone. What people actually hit is dilution, where every clause is present and each one gets less weight than it would have in a shorter prompt. The tell is that a dropped constraint from the middle reappears when you delete unrelated clauses elsewhere, which truncation cannot explain.

    Does the 50–150 word target apply to video prompts too?

    It is a reasonable starting band, but the per-family spread is much wider on video than on image. Wan's own guidance is a plain sentence plus a style word; Kling, VEO and Sora all reward dense camera language. Writing a Kling-shaped prompt for Wan over-engineers it against a schema that has nowhere to put the detail, and writing a Wan-shaped prompt for VEO under-specifies a model that was tuned to expect named moves.

    What if I need more than three hard constraints?

    Stop writing and change the input type. Three is roughly where prose stops being a reliable carrier, and past it you want pixels: reference images for identity and product, a locked first frame for composition, a schema field for anything the model exposes one for. A fourth constraint in text costs you one of the first three far more often than it adds a fourth.

    Should I write prompts as comma lists or full sentences?

    Follow the family. Midjourney's own guidance is comma-separated descriptors with filler words removed. Flux and Imagen both ask for natural language. The two styles are not interchangeable: a comma list on a natural-language family reads as a set of weakly related tags, and a paragraph on Midjourney spends most of its length on connective words that carry nothing.