Guides

    Night scenes that don't turn to mud

    Low-light prompts come back grey because nothing in frame motivates the exposure. The one-hard-source-plus-one-practical recipe and the post pass order.

    Versely Team8 min read

    Ask for a night scene and you will usually get a day scene with the brightness pulled down. Everything is visible, nothing is bright, the shadows are a flat charcoal with no detail in them, and the whole frame sits in a narrow band of grey that no amount of grading will rescue. That is mud, and it is not a rendering limitation. It is what happens when a prompt asks for darkness without giving the model anything that explains the exposure.

    Real night footage is not uniformly dark. It is mostly dark with a few things that are genuinely bright, and the contrast between those two states is what your eye reads as "night". A prompt that only says "at night" specifies the first half and leaves the second half to the training-set average, which is even, sourceless fill. The fix is to put something in the frame that has to be bright.

    Why "night" alone gives you grey

    Mood adjectives operate on the grade, not the geometry. dark, moody, atmospheric, noir, low-key all bias output toward desaturation and higher contrast without moving a single light. Stack five of them and you get five nudges to the same slider, which is how you end up with a frame that is simultaneously dim and flat.

    There is also a structural bias working against you. Models are trained on images that are, overwhelmingly, correctly exposed. Underexposure is rare in the data and usually accidental, so the prior pulls hard toward "everything readable". A prompt has to actively fight that, and adjectives are not enough force.

    The counter is motivation: a named light source that exists inside the frame, with a position, so the model has a physical reason for some pixels to be bright and the rest not to be. The vocabulary that works here all names something that emits — practicals, motivated light, the green glow of a monitor, harsh fluorescent overheads, volumetric, god rays. They land for the same reason "camera left" lands: they are executable.

    The recipe: one hard source, one practical, deliberate falloff

    Three clauses. Do not add a fourth until the first three are landing.

    1. One hard key. Small source, hard shadow edge. Streetlight, headlamp, a phone screen, moonlight through a window. Hard is the point — a small source at night is what gives you a shadow with a defined edge, and a defined edge is what stops the frame reading as fog. Say where it is relative to the subject and say it is hard: a single sodium streetlight high and behind her at camera right, hard-edged shadows.

    2. One practical in frame. Something visibly emitting, inside the shot, that the viewer can see. A neon sign, a lamp, a dashboard, a laptop. This is doing two jobs: it justifies the ambient level, and it gives the eye a bright anchor so everything else can be dark by comparison. Keep it to one. Two practicals in different colours is a look; three is a mess, and at that point you are back to even illumination.

    3. Falloff, stated. This is the clause almost everyone skips, and it is the one that decides whether the frame reads as night. Falloff is what happens to everything the key does not reach. Say it explicitly: the far side of the room falls to near-black with no detail, or light drops off within two metres of the lamp, the corridor beyond is unlit.

    Assembled:

    Medium shot, a woman standing at a rain-wet kerb at night. A single sodium
    streetlight high and behind her at camera right throws a hard-edged shadow
    across the pavement. A neon pharmacy sign at frame left is the only other
    source, cold green, reflecting in the puddles. The street beyond the
    streetlight's pool falls to near-black with no detail. Warm-cool split
    between the two sources.
    

    Nothing in that prompt says "dark", "moody" or "cinematic". It is four physical facts, and it produces night because night is the only arrangement consistent with them. The general source-direction-quality structure underneath this is covered in lighting prompts for AI video and images; this is that formula pushed to its low-light extreme.

    Say what you want, not what you don't

    The reflex for a muddy render is to reach for a negative promptno grey, no washed out, no flat lighting. On flow-matching models, which run without a classifier-free-guidance channel to push against, that text is close to a no-op, and even where negatives still function, instructive phrasing leaks the negated noun back into the frame. Google's guidance makes the same point with its own example: do not write "no buildings", write "a desolate landscape with no buildings or roads". The full case, and what replaced the field, is in negative prompts for AI video and images.

    For night work the practical move is translating every complaint into a positive:

    What you want to remove What to write instead
    "not washed out" "the shadow side falls to near-black with no detail"
    "no flat lighting" "hard-edged shadows from a single small source"
    "not grey" "deep saturated blue in the shadows, warm sodium in the highlights"
    "no daylight look" "the only illumination is the streetlight and the sign"
    "no visible ambient fill" "unlit corridor beyond the doorway"

    The right-hand column also happens to be better prompt material in general, because each line adds information rather than removing it.

    The post pass, in the order that matters

    Night footage is where generation artefacts are most visible, because compression and upscaling both behave badly on large dark areas. Order the finishing passes correctly or you will amplify what you meant to remove.

    1. Clean and deflicker first. Shadow noise and frame-to-frame flicker are cheapest to remove before anything else touches the pixels. Skip this and every later pass treats the noise as detail worth preserving.
    2. Then 1080p. Get to a solid intermediate before you reach for a final resolution.
    3. Then step up to 4K. Going straight from a raw generation to 4K amplifies artefacts you could have removed for far less effort at the start. Upscaling to 4K is a finishing move, not a fixing move.

    One more upstream lever: frame rate. Very high rates starve the model's temporal context and make flicker worse, so generating in the moderate band rather than at 60 is the cheaper prevention — 24 to 30 holds together, and temporal consistency is the property you are protecting. Versely's timeline default is 25 fps, so whatever rate a clip is generated at gets conformed on the way in; that conform is far kinder to a clean 24 or 30 than to a 60 that is already shimmering.

    A word on reviewing night work: the editor's 480p preview pass is free and carries a short per-user cooldown, which makes it excellent for checking composition, timing and whether a source is in the right place. It is not the right surface for judging shadow detail or banding, because both are exactly what a low-resolution preview destroys. Use the preview loop for structure, then judge the grade at full resolution once.

    FAQ

    Why does adding more light sources make it worse?

    Because each additional source raises the floor. Night reads as night because of the ratio between the brightest thing in frame and the darkest, and every source you add lifts the darkest. Two sources with a large gap between them beat four sources at similar levels every time. If a scene needs to be busier, add reflections and wet surfaces rather than more emitters — they multiply the apparent complexity of the light without raising the ambient level.

    Should I generate bright and darken in post instead?

    No. You cannot recover a shadow that was never dark, and darkening a correctly exposed render just compresses everything toward the middle, which is the mud you were trying to avoid. Grading can shift a look; it cannot invent a contrast ratio that was not generated. Fix it in the prompt.

    What about night scenes with dialogue where the face has to be readable?

    Add a small, close, soft fill and say that it is small and close: a soft fill from a phone screen just below her face, close and weak. That is a real thing a person can be lit by at night, so it satisfies the model's need for a motivation, and being close and weak means it lights the face without lifting the background. Naming a plausible in-frame reason is what keeps it from becoming generic ambient fill.

    Does this work the same for images and video?

    The prompt recipe does. The post pass is video-specific, and video adds one failure the stills workflow does not have: a source that changes brightness or position between frames because nothing pinned it. If a night clip flickers in a way that looks like the light itself is moving, that is usually the source being re-interpreted frame to frame, and the fix is to establish the shot as a still first and animate from it rather than generating the whole thing from text.