Scale cues that make a subject read as huge
Models default to mid-scale unless something anchors the frame. The lens, camera height and foreground cues that make a subject read enormous.
Prompt "a giant robot standing in a city street" and you will usually get a robot about the size of a bus, framed dead centre, with the camera at the robot's chest height. Every element of that image is arguing against the word you wrote. The camera is at a height that only makes sense if the photographer is floating twenty metres up, and there is nothing in frame whose real size the viewer already knows.
Scale is not an adjective. It is a set of relationships between the subject, the camera and something the viewer can measure against. Get those relationships into the prompt and you never need the word "giant" at all. Leave them out and no amount of emphasis fixes it, because the model has no size information to act on and falls back to the most photographed framing of whatever noun you used.
Why "giant" resolves to mid-scale
Two things are happening. The first is that size adjectives are labels, and labels resolve to the middle of a category the same way emotion words do. The second is more specific to images: almost every photograph a model has seen of a "robot in a street" was taken at human eye height with the subject fully inside the frame, because that is how photographs of objects work. The prior is not neutral — it is actively pulling toward the framing that made the subject fit.
So the job is not to persuade the model that the subject is large. It is to describe a photograph that could only have been taken of something large.
The cues that actually carry scale
A known-size object in frame. This is the strongest single cue and the one most prompts omit. A human figure, a car, a doorway, a streetlight, a pigeon. The viewer measures the subject against the reference, not against your adjective. Put the reference low in the frame and small, and specify where: "a single human figure at the base of the left leg, reaching roughly to the knee joint."
Camera height and angle. A low angle, and specifically a worm's-eye view with the lens near ground level looking steeply up, is what a photograph of something enormous looks like. It also produces converging verticals, which the eye reads as "very tall" before it reads anything else. High angle does the opposite and will fight you. Say the height explicitly rather than just the angle: "camera at ground level, tilted steeply upward."
Lens choice. A wide-angle focal length close to a subject exaggerates the difference in size between near and far parts of it, so the near foot is enormous and the distant head is small. That is the visual signature of standing next to something you cannot fit in frame. Focal length is not a vibe word in current models; wide-angle, portrait-length and macro all shift field of view and depth of field in a way you can see when you A/B them.
Foreground occlusion. Something close to the lens, cut off by the frame edge, slightly out of focus, and partially covering the subject. It establishes a near plane, and once there are two planes the viewer can infer distance, and once there is distance the subject's angular size means something. A branch, a railing, the back of a head, the corner of a rooftop.
Atmospheric haze on the far parts. Distant things lose contrast and shift toward the colour of the air. If the top of your subject is hazier than its base, the viewer concludes the top is much further away — which is only possible if the subject is very large. This is the cue that sells scale with no reference object at all.
Partial framing. Let the subject exceed the frame. A subject fully contained in the image is, by definition, something that fits.
| Cue | What it tells the viewer | Prompt phrasing |
|---|---|---|
| Known-size reference | Direct measurement | "a person at the base, reaching the ankle" |
| Ground-level camera, steep tilt up | Converging verticals | "camera at ground level, tilted steeply upward" |
| Wide-angle lens, close | Near/far size disparity | "24mm wide-angle, close to the near foot" |
| Foreground occlusion | Two depth planes exist | "out-of-focus railing across the lower third" |
| Atmospheric haze | The far end is genuinely far | "the upper structure fading into haze" |
| Subject exceeds frame | It does not fit | "the head cropped above the frame line" |
Three of those in one prompt is usually enough. All six at once tends to produce a mess, because haze and heavy foreground occlusion compete for the same part of the frame.
Order the prompt so the cues survive
The shot-card order most video prompt guides converge on puts cinematography first and style last: camera, then subject, then action, then context, then look. Google's Veo 3.1 guidance uses that skeleton, and the reason it matters here is that framing instructions placed at the end of a long prompt get dropped disproportionately often. If your lens and camera height are in the last clause, expect them to be treated as optional.
So this:
Worm's-eye view from ground level, 24mm wide-angle, tilted steeply upward.
A weathered steel structure fills the frame and continues past the top edge.
A single figure in a yellow jacket stands at its base, reaching roughly to the
first joint. An out-of-focus chain-link fence crosses the lower third.
The upper structure fades into pale morning haze. Overcast daylight, flat and
even.
Not "a giant steel structure, cinematic, dramatic, epic scale."
One trap to avoid while stacking cues: contradictory pairs. "Low angle bird's-eye view" or "wide-angle macro" will not average into something in between. The model picks one, apparently at random, and you will spend three generations wondering why the angle keeps changing. Every cue has to be physically consistent with every other one, because you are describing a photograph a real camera could have taken.
Scale cues are cheap to test, because the prompt changes are small and the difference is obvious at thumbnail size. Run the same subject four ways on a text-to-image generation — no cues, reference object only, camera and lens only, then everything together — and compare. A model with strong prompt adherence on layout, like Seedream 5.0 Pro, makes the comparison cleaner. Shot composition prompts covers the rest of the framing vocabulary.
Making a subject read as tiny
Every cue inverts cleanly, which is a useful check that you understood why they work.
- High angle, camera above rather than below. Bird's-eye at the extreme.
- Longer lens or macro, which compresses rather than exaggerates near/far difference.
- Very shallow depth of field, the plane of focus only centimetres deep. This is the basis of the tilt-shift look that makes real cities read as models: the eye associates that little depth with being very close to something very small.
- Oversized surroundings. A blade of grass that towers, a coin, a fingertip entering frame.
- Subject comfortably inside the frame, with margin on all sides.
Isometric diorama prompts push this further, into territory where the miniature reading is the point of the image rather than an effect applied to a realistic scene.
Video adds parallax, and takes away some control
Motion gives you two cues a still cannot have. Parallax is the strong one: as the camera moves laterally, near objects sweep across the frame quickly and far objects barely shift. A slow tracking move past a foreground element does more for perceived size than any static composition, because it demonstrates the depth rather than implying it.
The second is travel time. A camera that rises for four seconds and still has not reached the top of the subject has said something no still frame can say.
The complication is repeatability. Speed and easing are re-estimated every generation, so "slow crane up" is a different move each take and matching two shots becomes guesswork. Where you need the same move twice, camera control as a parameter beats camera control as prose, and describing the camera versus giving it a reference clip covers that trade. Framing the move as a first-and-last-frame transition pins both endpoints instead of one.
Remember too that models add movement by default. If the shot is meant to be locked off so the subject's stillness reads as mass, say the camera is entirely motionless. Silence gets you a drift.
FAQ
Does adding "giant" or "colossal" hurt?
Not exactly, but it does very little on its own and it can cost you. Size adjectives sometimes pull the whole image toward fantasy-illustration styling, which is wrong if the goal is a photographic result. If the physical cues are in place, the adjective is redundant. If they are not, it will not save you.
Why does the reference figure come out the wrong size?
Usually because the prompt named the reference without placing it. "A person for scale" leaves the model free to decide the ratio, and it will decide conservatively. Anchor it against a specific part of the subject: "reaching roughly to the ankle," "about as tall as one window." Giving it a landmark to match makes the ratio checkable rather than invented.
Can I fix scale after the fact?
Rarely, and not by grading or upscaling. Scale lives in perspective and framing, both of which are baked into the geometry of the render. A crop can help a little by removing frame margin, which makes the subject read as less contained. Anything more than that means regenerating with the cues in the prompt.
How many cues before it looks fake?
The failure is not too many cues, it is contradictory ones. A ground-level wide-angle shot with haze and a foreground element is exactly what a real photograph of a large object looks like. What reads as fake is a low angle combined with compression from a long lens, or heavy haze with hard corner-to-corner sharpness, because no single lens produces both.