Prompt pairs that cancel each other out
Contradictions like 'static handheld' make a model pick one instruction at random, take to take. The common conflicts and how to collapse each into one.
"Static handheld shot" is two instructions, and they are the opposite of each other. Handheld means an operator's body is holding the camera and their breathing is in the frame. Static means nothing moves. No camera can do both, so the model does what it always does with an impossible request: it satisfies one and quietly drops the other. Which one it drops is not something you control, and it is not the same on the next take.
This is a distinct failure from a model ignoring an instruction. A dropped instruction gives you a consistent miss — the move never happens, take after take. A contradiction gives you a bimodal result: some takes are locked and some are shaky, both technically correct readings of what you wrote. That inconsistency is why contradictions get misdiagnosed as an unreliable model rather than a broken prompt.
The tell: two families of output, not one
Run four takes of a prompt with different seeds and look at the spread. Normal variation looks like four versions of the same shot with different details. A contradiction looks like two versions of shot A and two versions of shot B, with nothing in between.
That is the diagnostic, and it separates two problems that feel identical from inside a single generation. Hold the seed fixed and you get the same resolution every time, which hides the issue entirely: a contradiction only shows itself across seeds. If your takes cluster into two groups, stop rerolling and read the prompt.
Where contradictions come from
Almost nobody writes "static handheld" on purpose. They arrive three ways.
Prompt accretion. A take comes back wrong, so you add a word. It comes back wrong differently, so you add another. Nothing gets deleted. After five rerolls the prompt holds the residue of five theories about what went wrong, and at least two disagree. This is by far the most common source, and the fix is to rewrite from scratch rather than patch, at least every third reroll.
Copied fragments. Prompt libraries, template lines and boilerplate style suffixes get pasted together. A style block ending in "handheld documentary feel" is a camera instruction hiding in the style slot, and it will fight the tripod you specified up front. Style presets do this invisibly, carrying camera and lighting assumptions that never appear in your text at all.
Wanting both things. Sometimes the contradiction is honest. "Static handheld" usually means "mostly still, but with a little organic life in it," which is a real and reasonable request. It just is not two instructions. It is one instruction with a magnitude.
The common conflicts, and what to collapse each into
| Conflicting pair | Why they fight | Collapse to |
|---|---|---|
static + handheld |
Handheld motion is the operator's body; static is no motion at all | Locked-off frame with a barely perceptible drift, as if on a tripod with a loose head |
locked-off + tracking shot |
Tracking requires the camera to travel with the subject | Locked-off wide; she walks through the frame and exits |
dolly in + zoom in |
Two different mechanisms for the same apparent result, at unspecified rates | Name one mechanism: the camera physically travels forward, focal length unchanged |
drone shot + no camera movement |
Drone implies flight | Static aerial frame from a hovering drone |
24mm wide-angle + heavy bokeh |
Short focal lengths produce deep apparent focus | Longer lens, or accept the deeper focus |
shallow depth of field + everything sharp |
Opposite focus states | Pick the plane, name both sides |
golden hour + harsh overhead fluorescents |
Two light sources with opposite size, direction and colour | Pick the motivated source and name its position |
macro + wide establishing shot |
Opposite framing scales | Two shots, not one |
slow motion + fast-paced, energetic |
Speed conflict | Keep the slow motion; carry the energy in the subject's action |
photorealistic + stylized illustration |
Opposite render modes | Pick one and put the other in a separate generation |
minimal, clean, empty + richly detailed, layered |
Opposite densities | Pick one, then specify what fills or empties the frame |
high angle + looking up at her |
Geometrically impossible | Pick the camera's height and stay in it |
symmetrical centred composition + rule of thirds |
Two placement rules for the same subject | Pick one |
8 seconds + a three-beat narrative |
More events than the duration holds | One beat per clip, or a multi-shot approach |
The right-hand column is the part worth studying. Very few of these get fixed by deletion. They get fixed by turning two competing instructions into one instruction that has a size.
"Locked-off frame with a barely perceptible drift" is a single coherent thing an operator could execute. So is "handheld with restrained natural sway, no whip movements." Both deliver what "static handheld" was reaching for without asking the model to choose. That reframing, from a pair of opposites to a point on a scale, resolves most stylistic conflicts too: minimal but warm beats minimal, clean, empty, richly textured.
The four categories
Sorting a conflict by type tells you how to collapse it.
Mechanical. Two physical camera behaviours that cannot coexist: static versus handheld, locked versus tracking, dolly versus zoom. These collapse by picking the mechanism and, where you wanted a hint of the other, adding a magnitude word. Never a second mechanism.
Optical. Two lens or light states: wide-angle against heavy background blur, golden hour against fluorescents, shallow against deep focus. These collapse by respecting the physics. Short lenses have deep focus; you do not get to override that by asking harder. Pick the state that matches the rest of the shot.
Stylistic. Two render modes or densities: photoreal against illustrated, minimal against maximal, warm against cold grade. These collapse by choosing one and then describing what you actually wanted from the other in concrete terms. If "minimal" and "detailed" both went in because you wanted a clean frame with one interesting texture, write that.
Temporal. Not two instructions fighting, but one instruction fighting the clock: three narrative beats in an eight-second clip, or "a full day passes" in a single generation. These collapse by splitting into shots, or by naming a device that legitimately compresses time, like a locked-off timelapse.
A sixty-second prompt audit
Before you spend another generation on a prompt that keeps coming back two different ways:
- Rewrite the prompt as a bare list. One instruction per line, imperative, no connecting prose. A 60-word paragraph usually decomposes into eight to twelve lines.
- Read the camera lines together. Pull every line that mentions the camera, the lens, movement or framing into one group, wherever it appeared in the original. Conflicts hide in the gap between the camera slot at the front and a style suffix at the end.
- Ask the operator question on each pair. Could one person with one camera execute both of these at the same time? If the answer needs a clarifying question, it is a conflict.
- Collapse, do not delete. For each conflict, write the single instruction with a magnitude that gets you what you wanted from both.
- Rebuild in slot order. Camera first, then subject, action, context, style. Re-reading in that order surfaces the leftovers, because a camera word sitting in the style slot is obvious once the slots are separated.
Then rerun four takes across seeds and check the spread again. If the two families collapsed into one, the prompt was the problem. If they did not, the problem is elsewhere and the wider diagnostic tree for generations that come back wrong is the next step.
One thing this is not
A contradiction is not the same as an instruction the model cannot execute at all. Some camera language is prose the model approximates no matter how cleanly you write it, a separate problem covered in camera instructions models silently ignore. The fixes differ: a contradiction is fixed by editing, an unsupported instruction by changing model or technique. Checking which video models expose real camera control as a parameter, in the model catalog, tells you which one you have.
Resolve conflicts before you run any model comparison. The same contradictory prompt across four models produces four different resolutions and no usable signal, which is a common way an otherwise sound model A/B test ends up inconclusive.
FAQ
Will the model tell me if my prompt contradicts itself?
No. There is no validation step and no warning. The generation succeeds, returns a well-formed clip, and looks fine until you notice it does not match what you asked for. Silent resolution is the whole problem — a contradiction that errored out would be a much easier bug to have.
Can I put one side of the conflict in the negative prompt instead?
Rarely a good idea. On architectures that run without a usable negative-guidance channel the field is effectively ignored, and even where it works, a negative prompt full of concepts you also mentioned positively is more confusion, not less. Resolve the conflict in the positive prompt where you can see it.
How long can a prompt be before conflicts become likely?
Length is not the variable; edit history is. A carefully written 120-word shot card can be entirely consistent, while a 30-word prompt patched four times often is not. The signal is how many times you have added a word without removing one.
Does the agent help with this?
It helps most on the comparison side. Because the agent can fan one prompt across several named models in a single request, you can see whether a conflict resolves the same way everywhere or differently per model — which is useful evidence, but only after you have already cleaned the prompt. Fan out a contradictory prompt and you are just collecting more coin flips.