AI Models

    Which video models hold up with crowds in frame

    Background people break more shots than hero subjects. The pixel budget that predicts failure, plus staging tactics for street, stadium and venue shots.

    Versely Team9 min read

    A hero subject gets the model's attention. Six people twenty feet behind them do not, and that is where generated video gives itself away. The subject's face is fine. The three pedestrians crossing behind her have swapped jackets between second two and second three, one of them has a second left arm, and the man at the back of frame is the same man twice, standing four feet apart.

    Crowds are the hardest thing you can ask a video model to do, and unlike hands or teeth it is not a texture problem. It is a budget problem. Understanding what the budget is makes it possible to predict which shots will work before you spend anything generating them.

    The three ways a crowd fails

    They are distinct failures with distinct fixes, and lumping them together is why people conclude "crowds don't work" instead of "this crowd doesn't work."

    Duplication. The same face, the same coat, the same gait appears twice or three times in one frame. This is the model reusing a learned pattern because it does not have enough signal to individuate. It shows up most in mid-ground clusters where people are large enough to have features but small enough not to warrant them.

    Edge dissolution. Limbs at the perimeter of a group melt into each other or into the background. Arms merge where two people overlap. A leg passing behind another body does not come out the other side. This is the model failing at occlusion ordering, which is the same underlying weakness that makes hands difficult, applied to a scene with many more overlapping objects.

    Reset. The crowd is coherent in second one and coherent in second four, but it is not the same crowd. Identities do not persist; the background population is re-rolled as the clip proceeds. This is a temporal consistency failure specific to non-hero subjects, and it is the one that survives a single-frame review and kills the shot in playback.

    Pixel budget is the number that predicts it

    Before you argue about models, do the arithmetic on your shot. It takes thirty seconds and it is more predictive than any leaderboard.

    A background person standing at what reads as middle distance typically occupies about one twelfth of frame height. At a 720p ceiling (1280 x 720), that is 60 pixels of person. A head is roughly one seventh of a standing figure, so about 9 pixels of head, and a face is a fraction of that. At 1080p the same figure is 90 pixels and the head is about 13.

    Nine pixels is not a face. It is a smudge with a hairline, and a model cannot hold identity on a smudge because there is nothing there to hold. Thirteen is not much better, but it is the difference between reading as a person and reading as a glitch.

    That gives you a usable rule. If a background figure is under about 10 percent of frame height, do not expect it to survive as an individual — expect it to survive as a shape. Design the shot so that shapes are all you need from it.

    The second half of the budget is time. Versely renders at 25 frames per second by default, not 24, so an 8-second clip is 200 frames. Every one of those frames is an independent opportunity for a background identity to drift. Six background people over 200 frames is 1,200 identity decisions in one generation. That is why crowd failure scales with duration far faster than hero-subject failure does.

    Which model classes cope

    The catalog splits into four classes on properties that are visible before you generate anything, and those properties map directly onto the two budget terms above.

    Class Example Resolution ceiling Duration window Crowd implication
    Short-window, high ceiling Veo 3.1 up to 4K 4s / 6s / 8s enum Fewest frames for drift, most pixels per background face
    Long-window flagship Seedance 2.0 1080p 4s to 15s More seconds means more reset opportunities
    Fast and distilled tiers Sora 2 text-to-video 720p 4s to 20s in 4s steps Smallest pixel budget per background figure
    Conditioned on an input Happy Horse 1.1 image-to-video 1080p 3s to 15s Crowd layout is fixed on frame one, not invented

    The fourth row is the one worth taking seriously. When you supply a still image, the model is not deciding how many people are in the shot or where they stand. It inherits that from the plate and only has to move it. This does not eliminate reset, but it removes the entire class of "the model invented seven strangers with nothing to base them on" failure. If your crowd matters, image-to-video and reference-to-video paths are structurally better positioned than text-to-video, regardless of which family you prefer.

    The pattern across classes is simple: crowds reward short clips at high resolution and punish long clips at low resolution. A 6-second 1080p generation with four background figures is a reasonable ask. The same crowd at 720p over 15 seconds is not, and no amount of prompt work fixes it, because the pixels the model would need are not being rendered.

    Staging: street, stadium, venue

    Three common briefs, three different tactics.

    Street. The instinct is to write "busy city street, lots of pedestrians." That maximises the number of mid-ground figures, which is exactly the region where duplication lives. Invert it. Push the crowd either much closer (one or two people in the foreground, deliberately partial, cropped by the frame edge) or much further (silhouettes against a bright background, small enough that shape is all you were ever going to get). The middle distance is the danger zone; empty it.

    Foreground occluders help twice. A parked car, a railing or a passing bus across the near third of frame gives the model a hard object to compose against, and hides the region where limb merging happens.

    Stadium. Never ask for individuated spectators. A stadium crowd in real footage is texture, not people, and that is good news: the model is being asked to produce a field of colour with coherent motion rather than 4,000 identities. Keep the crowd behind a foreground subject, keep it slightly out of focus, and describe it as a mass with a single collective behaviour ("a wall of spectators rising to their feet in unison") rather than as a set of individuals.

    Venue. Interior events are the hardest of the three because lighting is low, bodies overlap heavily, and the camera is usually close enough that faces land in the 10-to-20-percent-of-frame-height band where the model tries to individuate and fails. The fix is aggressive depth separation: one lit subject, everyone else in shadow or backlit. Silhouettes at a venue read as atmosphere, not as evasion.

    Across all three, use the negative prompt field where the model exposes one rather than writing "no extra limbs" into the scene description. Kling, Pixverse, LTX, Wan and Veo's fal variant all take exclusions as a separate parameter that is read independently of the prompt, and that is where terms like extra limbs, duplicate faces and distorted background figures belong.

    Testing it without burning a day

    Crowd behaviour is not something you can read off a spec sheet, so run a small controlled comparison once and reuse the result for months.

    1. Write one prompt with an explicit background count. "Three pedestrians crossing behind her" is testable; "a busy street" is not, because you cannot score a failure against an unspecified target.
    2. Generate it on three or four candidate models. Versely's agent chat can fan a single prompt across several named models in one request, which is what makes this practical rather than tedious.
    3. Score two things only: did the count survive, and did the identities survive. Everything else is a different test.
    4. Re-run the winner at two more seeds. A model that gets a crowd right once and wrong twice is not a model that gets crowds right.
    5. Check the result in playback, not in stills. Reset is invisible frame by frame.

    For assembly, the editor's 480p preview pass is free and carries a short per-user cooldown, so you can check pacing before committing. The final export is charged once regardless of clip count, so a crowd sequence built from six short generations does not cost six exports. Reviewing a video before you publish it is the step where reset usually gets caught.

    FAQ

    Is there a model that just does crowds properly?

    No, and any post claiming otherwise is describing one lucky generation. What exists is a set of conditions under which crowds work across families: short duration, high resolution ceiling, background figures either large and few or small and shapeless, and a conditioning image where possible. Choose a model that lets you hit those conditions rather than one with a reputation.

    Why do crowds look worse in my clip than in the model's own demo reel?

    Demo reels are staged around the constraint. Look closely at published crowd shots from any provider and you will usually find heavy depth-of-field separation, silhouettes, motion blur, or a crowd that is genuinely just texture. Those are the same tactics described above, applied by people who know where the failure lives.

    Does upscaling a 720p crowd clip fix the faces?

    No. Upscaling increases pixel count, it does not recover identity that was never generated. A 9-pixel head upscaled to 4K is a large smudge. Generate at the resolution you need the detail at, then upscale for delivery if you want, not to rescue detail.

    Can I fix a duplicated face in post instead of regenerating?

    Sometimes. If the duplicate is small and static, an inpaint or a targeted patch over the region can work. If it moves across frames, you are doing per-frame work and regenerating with a foreground occluder in the shot design is almost always faster.