Fluid, fire and smoke: where video models diverge
Pours, flame and volumetric smoke fail in three different ways. How to rank them, and a rule for what to generate, what to composite, and what to cut.
There is a category of shot where "does it look good" is the wrong question. A pour of whisky into a glass either behaves like a liquid or it does not, and if it does not, no amount of grade or grain rescues it. The viewer will not be able to name what is wrong. They will just distrust the shot.
Fire, smoke and liquid are the three materials where this happens most, and they are not one problem. They fail for different reasons, at different rates, and the right response to each is different. Ranking them properly is the difference between a shot list you can shoot and one that burns a week.
Rank them by how much the material forgives
The useful ordering is not "which looks hardest" but "how much error the material's real-world behaviour hides."
Fire is the most forgiving. Real flame is chaotic and high-frequency: at 25 frames per second, the default rate here rather than 24, a flame changes shape substantially between adjacent frames anyway. A model producing temporally incoherent flame is producing something a viewer already expects to be incoherent. Fire is the one material where the model's weakness and the material's nature align.
Its real failure mode is elsewhere. Flame is self-luminous, so it should be lighting the scene around it. The tell is not the fire, it is everything next to it: a candle burning with a perfectly steady face beside it, a bonfire casting no moving light on the ground, a match struck in a dark room where the room stays dark. Models render flame far more reliably than they render its consequences.
Smoke is in the middle. Soft edges forgive a great deal per frame, which is why a still from a smoke shot almost always looks fine. What smoke fails at is accumulation. Real smoke has a history: it builds in a room, drifts one way, thickens where it is trapped. Generated smoke often behaves as a stationary texture that neither accumulates nor clears, or dissipates and re-forms with no memory of where it was. It also fails at depth ordering, sitting in front of objects it should be behind.
Liquid is the hardest, by a wide margin. It combines three things models are individually weak at. Volume has to be conserved, so the level in the receiving glass must rise monotonically as the stream pours. The stream is a continuous surface that must stay attached to the vessel at one end and the receiving surface at the other. And liquid is specular, carrying reflections of a scene the model also has to keep consistent. A pour is not one hard thing, it is three hard things that must agree. The general weight-and-contact problem is covered in physics failures in generated motion; liquid is that problem plus optics.
The specific tells, so you can review fast
Reviewing this material is a checklist, not a vibe. Watch these and nothing else on the first pass.
| Material | Watch this | Failure looks like |
|---|---|---|
| Pour | The receiving vessel's fill level | Level stays flat, or jumps, or the glass is full before the stream stops |
| Pour | The two contact points | Stream detaches from the bottle lip, or lands without disturbing the surface |
| Splash | Droplet count over time | Droplets appear from nothing mid-air, or a crown forms with no impact |
| Flame | The lit surfaces, not the flame | No flicker on adjacent faces, walls or ground |
| Flame | The base | Flame floats off the wick, log or burner rather than attaching to it |
| Smoke | Density over the clip | Neither builds nor clears; the room is equally smoky at 0s and 8s |
| Smoke | Occlusion order | Smoke drawn in front of a foreground object it should pass behind |
The fill-level check is the highest-yield item here. It takes four seconds, catches most bad pours, and is completely objective, which means you can hand it to someone else and get the same answer.
Where the model classes actually diverge
The catalog does not label models by fluid quality, and no honest post can tell you which one wins a pour, because the answer moves with every release. What you can read off the catalog are the structural properties that determine whether a fluid shot has a chance.
Duration window. Accumulation failures scale with time. A 4-second smoke shot rarely needs to demonstrate that smoke builds; a 15-second one does. Models with short fixed enums, like Veo 3.1 at 4, 6 or 8 seconds, are being asked a structurally easier question than models offering 15-second windows such as Seedance 2.0 or Wan 2.7 text-to-video. That is not a quality statement about either. It is a statement about what you are asking.
Native audio. Fire and water are two of the few subjects where synchronised sound does real work. A crackle that lands with a spark, or a splash that lands with the impact, buys plausibility the picture alone did not earn. Models with native audio are worth a look specifically for these materials.
A real negative prompt field. Kling, Pixverse, LTX, Wan and Veo's fal variant expose negative prompt as a separate parameter read independently of the scene description. For fluid work that field earns its keep: floating droplets, disconnected stream, static smoke are exclusions, and exclusions belong there rather than folded into the prompt as "no floating droplets," which the model reads as a request containing the words floating and droplets.
Conditioning. Starting from a still fixes the vessel, the liquid level at frame one, and the lighting, which removes an entire tier of invention from the model's job. For anything involving a product in a glass, an image-conditioned path beats pure text-to-video.
Generate, composite, or script out
This is the decision the whole article exists to support.
Generate it when the material is atmospheric rather than load-bearing: smoke drifting through a shaft of light, a fire in the background, steam off a cup. Nothing in the shot depends on the material behaving correctly, and soft-edge forgiveness works for you. Also generate when the clip is short and the camera is close on the flame itself, where fire's natural incoherence covers for the model.
Composite it when the material must interact with something specific and you already have that something. Generate a clean plate with no fluid in it, then bring a real element in over the top. This is where video segmentation and background removal do the work: the model produces the shot, the element produces the physics, the timeline joins them. The plate is the easy half, and it is the half models are good at.
Script it out when a pour, a splash or a fill is the hero action and has to be perfect. This sounds like defeat. In practice it is the highest-leverage decision on this list, because there is almost always a cut that gets the same story beat without the physics: the hand tilting the bottle, cut, the full glass being lifted. Two shots the model handles well, replacing one it handles badly. Audiences read that as competent editing, not avoidance. Nobody ever needed to see the whole pour. They needed to believe it happened.
Prompting that moves the needle
Four things that consistently help, and one that does not.
- Name the container and the receiving surface explicitly. "Amber liquid pours from a cut-glass decanter into a short tumbler on slate" gives the model three anchors. "A drink being poured" gives it none.
- Specify one direction of change. "The tumbler fills to halfway over the shot" tells the model the level must move. Without a stated direction, a flat level is a perfectly valid interpretation of the prompt.
- Describe the light the fire makes, not just the fire. "Firelight flickering across her face and the wall behind her" is prompting for the tell that usually fails, which is the point.
- Give smoke a destination. "Smoke drifting left across the frame and thinning as it goes" encodes accumulation as a visible change. "Smoky room" encodes a static texture.
What does not help: intensity adjectives. "Dramatic, powerful, explosive splash" moves the aesthetic register and does nothing for the physics. If volume was not conserved at "splash," it will not be conserved at "explosive splash." Where a model exposes a motion level parameter, that is the actual control, not the adjective stack.
When you assemble, the editor's 480p preview pass is free and carries a short per-user cooldown, which makes it cheap to check whether a composited element sits correctly in the cut. The final export is charged once no matter how many clips are on the timeline.
FAQ
Which model is best for water?
Nobody can answer that honestly for longer than a release cycle, and posts that do are describing a single generation. The durable answer: pick a model with a short duration window, a real negative prompt field and an image-conditioned entry point, then test your specific shot on two or three and keep the result.
Why does my fire look fine but the scene look wrong?
Because the model rendered the flame and skipped its consequences. Flame is self-luminous and should be driving the lighting of everything near it. Prompt for the lit surfaces specifically, or shoot the fire against a dark, mostly empty background where there is nothing for it to fail to light.
Is generated smoke ever good enough for client work?
Regularly, when it is atmosphere. Haze in a beam of light, breath in cold air, steam off food: nothing depends on accumulation and soft edges hide frame-to-frame drift. It stops being good enough the moment the smoke has to tell the story, such as a room filling during a fire safety explainer.
Can I fix a bad pour by slowing it down?
Usually not. Changing playback speed stretches existing frames, so a stream that detached from the bottle stays detached and is now on screen longer. Slowing helps judder, not physics. Regenerate with the container and fill direction stated, or cut around the pour.