Guides

    Name two foley events, and no more

    Native-audio models place a limited number of sound events per clip. Name two specific, countable sounds and you get both; name six and you get mush.

    Versely Team9 min read

    A prompt that lists six sounds usually comes back with two of them, and rarely the two you cared about. A prompt that lists two comes back with both, placed on the right frame, with a believable room underneath them that you never asked for.

    That is the whole technique. The rest of this is why it works and how to write it down.

    Sound events and the bed are two different requests

    Native audio models generate dialogue, effects and ambience in the same pass as the frames, which is the reason the audio lands in sync with on-screen action without any work from you. It is also the reason the audio has a budget, in the same way the picture does. How far a model goes — ambience only, ambience plus synced effects, or full dialogue — varies, and the native-audio model map sorts the catalog into those tiers.

    The budget applies unevenly, because the soundtrack splits into two categories that behave nothing alike:

    Discrete events. Countable, tied to something visible, with a clear onset. A cup meeting a saucer. A door latch. A single footstep on gravel. These have to be scheduled — placed on a specific frame, with the right transient — and scheduling is the expensive part.

    The bed. Continuous, uncountable, no onset. Room tone, distant traffic, rain, the low murmur of a café. Beds are cheap. A model that knows it is in a café generates café hum without being told, because the environment description already implies it.

    The mistake is asking for beds as if they were events, and events as if they were beds. "Add kitchen sounds" is a bed request with no bed described and no event named, so you get whatever generic clatter the model has lying around. "The knife taps twice on the board, and a pan sizzles behind her" is two events plus an implied kitchen, and the bed builds itself.

    Which brings us to how an event should be written. The unit of a good foley cue is material plus surface plus action, not a mood.

    What people write What the model can schedule
    "coffee shop sounds" "the espresso machine hisses as she turns away"
    "a clink" "a glass sets down on marble with a hard click"
    "footsteps" "one boot heel strikes wet concrete"
    "office ambience with typing" "a single keystroke, then the chair creaks"
    "car noises" "the door thunks shut, and the indicator ticks"

    Each right-hand cell names a material (glass, marble, boot leather, wet concrete), a surface it meets, and a countable action. That is enough information to synthesise a transient and place it. "Clink" is not — it is a description of a sound the model already has fifty versions of, with no way to choose between them.

    Two is the working ceiling for a short clip. It is not a hard limit anyone published, it is where the failure starts: at three or four named events the model begins dropping one, merging two into a single smear, or firing them all in the first second because it ran out of room to space them. Two events in a five-second clip have space to breathe and to sit on the frames where the action actually happens.

    The syntax, one idea per sentence

    The VEO family responds to a labelled structure, and the labels genuinely matter — they separate what would otherwise read as one run-on description:

    • Dialogue goes in quotes: She says, "It's still warm."
    • Effects get an SFX label: SFX: a glass sets down on marble with a hard click.
    • Ambience gets its own: Ambient noise: low café chatter, no music.

    Keep each in its own sentence. Stacked audio instructions inside a single sentence get parsed as one blurred request. The VEO 3.1 prompting notes cover the rest of the syntax. Models that do not take those labels still benefit from the same discipline — one audio idea per sentence, stated plainly, placed after the action — because what helps is the separation, not the label itself.

    A complete two-event prompt, in the shape that works:

    Medium close-up, static camera. A woman in an apron slides a glass of water across a marble counter and turns away. SFX: the glass sets down on marble with a hard click. SFX: the espresso machine hisses behind her. Ambient noise: low café chatter. No music.

    One shot line, one action, two named events, one bed, one exclusion. Nothing in that prompt is decorative.

    Note what is absent: no "cinematic sound design", no "immersive audio", no "rich soundscape". Those are mood adjectives, and they behave the way mood adjectives behave in the lighting half of a prompt — biasing the overall texture without specifying anything the model can execute.

    Short clips schedule sound better than long ones

    Three to five seconds produces sharper foley than eight, and the reason is arithmetic rather than quality. An eight-second clip has more frames to fill and more implied events to invent, so the model spreads its attention across a longer schedule and the individual transients get softer. A four-second clip has one or two things happening, and it commits to them.

    This is convenient, because 3–5 seconds is also the band where temporal consistency holds up best, and it is roughly the cut length short-form platforms want anyway. The constraint that keeps faces from drifting is the same one that keeps foley crisp.

    The practical version: when a scene needs four sound events, it needs two clips, not one longer clip. Generate two four-second shots with two events each, then join them in the video editor. You end up with four clean events and a cut, instead of four soft events and no cut.

    Exclusions, and the one thing to never prompt

    Two events plus one bed leaves room for one exclusion, and there is almost always one worth spending it on. Reverb is the usual candidate — it arrives uninvited in a large proportion of interior generations and it is the single hardest thing to remove after the fact.

    Phrase exclusions as a described state, not an instruction. "No reverb" reads as an instruction containing the word reverb, and instructive negation has a habit of leaking the negated thing back into the output. "A small, acoustically dead room" describes a state the model can render. Same request, opposite failure rate. This is the same principle that makes negative prompts unreliable on modern architectures: you are better off describing the thing you want than naming the thing you don't.

    The one thing to leave out entirely is music. Prompted background music is the weakest output of every native-audio stack, and worse, it is baked into the mix — you cannot pull it down under a voice later because it is not a separate track. Generate dialogue, effects and ambience only, then add a track in the edit where you control the balance. Building a bed from effects before you reach for music covers that ordering properly.

    Confirming both events landed

    Listen on a phone speaker before you listen on headphones. Small speakers roll off low frequencies hard, so a foley cue that carries its weight in the bottom end — a door thunk, a heavy footstep — can be perfectly audible in monitoring and effectively gone in the feed. If you can hear both named events on a phone, they are there.

    Do the check on a preview rather than a finished render. The editor is EDL-based, so the timeline is a description you re-render rather than a file you destroy, and preview: true gives you a 480p pass at no credit cost with a short per-user cooldown between them. Audio survives the downscale intact, so the cheap pass genuinely answers the question. The final export is charged once regardless of how many clips are on the timeline — the previews and export breakdown has the exact shape.

    If one of the two events is missing, re-roll. If both are missing, the problem is the phrasing rather than the seed: check that each event names a material and a surface, and that each sits in its own sentence.

    FAQ

    Why does naming more sounds make the output worse?

    Because discrete sound events have to be placed on specific frames, and the model is dividing a fixed amount of scheduling attention across everything you named. Two events in a short clip get spaced properly. Five compete, and the usual result is that two survive, two vanish, and one smears into the bed. Ambience does not compete the same way, so a described environment costs you nothing.

    Should I name the sound or the action that makes it?

    Both, in one clause. "A glass sets down on marble with a hard click" gives the model a visible action to sync against and a material pair to synthesise from. Naming only the sound leaves it floating, unattached to any frame. Naming only the action means the sound is inferred, which works for obvious cases and fails for anything with a specific texture you wanted.

    Which models are worth using this on?

    Any model that generates audio in the same pass as the picture — but the capability is not uniform, so check before spending prompt effort on foley. Some produce ambience and little else, which leaves a named event with nothing to schedule it. The catalog's audio-capable list filters to models whose feature tags name an audio capability; confirm the tier, because synced effects is the band this technique needs.

    Can I add the third and fourth sounds later?

    Yes, and usually you should. Generate the clip with its two native events, then layer additional effects as a separate track in the edit, where they have their own level control and can be moved a few frames without re-rolling the video. Native audio is for the sounds that must be locked to picture. Everything else belongs on the timeline.