9:16 keeps cutting your subject off at the waist
Aspect ratio is a composition input, not a crop. Tall frames need shot-type language that reaches the feet, plus recovery options short of a re-roll.
9:16 is not a crop you apply to a finished 16:9 picture. It is the frame the model composes inside. Ask for a person at 9:16 without naming a shot type, and you will get head-to-waist more often than head-to-toe, because that is the composition a tall frame invites.
The model is not failing to follow "a woman in a linen suit." It is filling a phone-shaped canvas the way most phone-shaped training images are filled: face in the upper third, torso in the middle, legs implied below the fold.
If you need the shoes, say so in shot language and generate at 9:16 from the start. Outpainting the missing band works when the body was slightly clipped. Inventing a lower half that was never there is a worse job.
The model composes for the frame it was given
Aspect ratio is a composition decision. The model places the subject for the rectangle it will have to fill. The same prompt at 16:9 and at 9:16 is not one picture cropped two ways. It is two pictures. Widescreen invites a standing figure with environment left and right. Portrait invites a bust, because a full-length body in a tall thin frame is a small body with a lot of floor and ceiling, and models are biased against "small subject, empty frame" unless you demand it.
That is why generating wide and cropping down is the expensive option. You throw away the sides, and you also throw away the composition, because the subject was placed for a wide frame. The result is the glossary's own pitfall: the subject sits in the middle third, and the vertical crop leaves you with headroom and a chin. Where each ratio wins is the placement map. This article is the generation problem inside 9:16 itself.
Video inherits the same prior. A text-to-video model given "a man walking through a market, 9:16" will often plant him from the belt up. You wanted gait. You got a moving waist crop. The fix is shot type, plus enough empty canvas that the feet have somewhere to exist.
Shot-type language that reaches the feet
Name the shot. Name the body landmarks that must be in frame. Name the space above the head and below the shoes. Put that block at the end of the prompt, where camera behaviour tends to stick.
A full-length still:
Full-length standing portrait of a woman in a linen
suit, head to toe, both shoes fully visible, generous
space above the hair and below the feet, subject small
enough in frame that nothing clips the edge, 9:16,
studio, even light, locked-off
A walking shot for vertical video:
Wide full-length, man walking toward camera through a
market aisle, entire figure visible, feet and market
floor in the bottom of frame, head not touching the
top edge, 9:16, static camera, tripod, no tilt down
A knee-up when you do not actually need shoes, but you do need hands:
Cowboy shot, mid-thigh up, both hands in frame, not
cropped at the waist, 9:16, subject centered, talking
to camera, locked-off
If you want a talking head, say that too, so the model stops guessing:
Medium close-up, chest up, face in the upper third,
9:16, talking to camera, locked-off, no tilt
Shot type is a composition class, not a vibe. Cowboy / American is mid-thigh, the usual recovery when full-length comes back as a tiny doll on a corridor of floor. Landmarks (both shoes fully visible, both hands in frame) are checkable; "full body" is not. Margin (subject small enough in frame) fights the prior that wants the face large. If the face is large in 9:16, the feet are out by arithmetic. Locked-off, no tilt down stops a video model from "finding" the face and cropping the legs mid-clip.
Write the desired state. Prefer "feet visible in the bottom of frame" over "no missing feet."
| You want | Phrase that usually gets it | Phrase that usually does not |
|---|---|---|
| Shoes in frame | full-length, head to toe, both shoes visible, subject small in frame |
full body, cinematic portrait |
| Hands, not shoes | cowboy shot, mid-thigh up, both hands in frame |
medium shot (often lands at the waist) |
| Talking head | medium close-up, chest up, face in the upper third |
leaving shot type unsaid on 9:16 |
| Walk cycle | entire figure visible, feet and floor in the bottom of frame, static camera |
walking, 9:16 |
Generate at 9:16. Do not generate 16:9 and crop. If the deliverable is a Reel or a Short, start in the video generator at 9:16. Reformatting later is a separate job.
Recovery without re-rolling
You missed. The still is good except the shoes are gone. Re-rolling throws away a face, a grade and a pose you already like. Three recoveries, in the order you should try them.
1. Outpaint the missing band, when the clip is small. If the frame cut at mid-shin or at the ball of the foot, the model has real anatomy to continue. Mask a new canvas below the original, describe only the missing band (continue the linen trousers to leather oxfords, same studio floor, same light, no new objects), and generate that strip. Outpainting is the right tool when you are extending a composition that already almost fits. Keep the generation window tight against the real pixels; a wide-open pad with a sliver of ankle is how you get a second pair of shoes. The failure modes (bent lines, duplicated hardware, colour drift) are covered in uncrop without warping. Stage the extension. Do not jump from a waist crop to a full-length in one pass.
2. Stop if the cut is at the waist. A waist crop has no legs to continue. The model will invent a lower body from a clothing edge and a prior about people. Invented legs are where this goes wrong, the same way an arm that already exits frame-left becomes a hallucination when you extend left. Re-roll with the full-length block above. You are not being precious. You are refusing a reconstruction job the tool is bad at.
3. Use a platform-native reframe only when the body is already complete. Midjourney's Pan and Custom Zoom (--zoom n --ar x:y) recover a missed framing without a full re-roll, provided the subject is fully inside the source. That is a Midjourney control, not a Versely one. On Versely the equivalent is an outpaint via generate_image_from_image with a mask on the new canvas, or, for video that is already full-length and just in the wrong rectangle, make the video vertical. Those tools reframe existing pixels. They will not grow legs. Generate the vertical as its own shot.
A campaign that has to live in several ratios is a different design problem, covered in outpainting one creative into every ratio. Do not borrow that workflow to rescue a generation that never included the feet. It assumes the subject is already fully inside the master.
Generate at the ratio you will publish
If the post is 9:16, the generation is 9:16. If you also need a 16:9 cut, generate a 16:9 shot with its own shot type, rather than stretching one composition across both.
Practical sequence:
- Decide the publish ratio before you write the prompt. Not after the contact sheet.
- Put shot type, landmarks and margin at the end of the prompt.
- Generate two or three takes at that ratio. If every take is a waist crop, make the figure smaller in language (
subject small in frame,lots of floor). - If one take is close and only slightly clipped, outpaint the clipped edge. If it is cut at the waist, re-roll. Do not outpaint a torso into a person.
- Finish captions, grade and overlays on the ratio you will post, so safe zones are real. Platform UI eats the bottom ~15 percent and the top ~10 percent of a 9:16. Feet you barely kept will vanish under the caption stack.
For video, one camera move per shot, stated at the end. A tilt that "finds" the face is how a full-length prompt becomes a waist crop at second two. Static, locked-off, tripod is the lock. If you want a push-in, push in on a figure that already has margin below the feet, so the move cannot eat the shoes.
FAQ
Why doesn't writing "full body" fix the waist crop?
Because "full body" is not a shot type. The model can satisfy it with a figure whose feet sit on the bottom edge, or with a torso it considers complete. Head-to-toe, both shoes visible, and subject small in frame are checkable constraints. Use those.
Can I outpaint the missing legs?
Only when some of the legs are already in frame. A shin or a shoe to continue is enough to try. A waist cut is not. Invented lower bodies from a belt line fail often enough that a re-roll with better shot language is cheaper than a cleanup.
Should I generate 16:9 and crop to 9:16?
No. The model composed for the wide frame, so the vertical crop is headroom and a chin, or a waist. Generate at 9:16. If you need both placements, generate both shots.
What if I want a talking head, not shoes?
Then ask for a medium close-up, chest up, face in the upper third. The waist crop is the unguided default, not a talking-head request. Naming the shot stops the model guessing, in both directions.