Spatial audio that pans with the camera
Some models move the sound when the camera moves, making camera language audible. When that earns the prompt work, and how to check it survives a phone speaker.
Write "the camera arcs slowly around her" into a prompt for a model with spatial native audio, and the café behind her moves too. The espresso machine that was on the left ends up on the right. Nothing in the prompt said anything about sound.
That is the whole feature, and it is easy to over-value. Spatial audio makes camera language audible, which is a real creative capability, and it is also inaudible to a large share of the people who will watch the result. Knowing which side of that line a video falls on is the entire decision.
What "pans with the camera" actually means
Native audio models generate sound in the same pass as the frames rather than as a layer added afterwards. Because sound and picture come out of the same process, a model that has learned spatial placement can keep a sound source anchored to its position in the scene while the frame moves around it.
Treat that as emergent behaviour rather than a switch you turn on. No model in the catalog publishes a spatial-audio specification, and the audio feature tags that exist say whether a model produces sound at all, not how it positions it. Whether a given model anchors a source to its place in the scene is something you establish by generating an orbit and listening.
The distinction worth holding onto: this is source-anchored audio, not stereo width. A wide stereo mix sounds big and stays put. Spatial audio keeps a specific sound attached to a specific object, so when the camera orbits, the object's sound orbits with it in the opposite direction — the same way it would if you walked around a real room.
The effect only exists when two things are true at once: the camera moves, and there is a sound source in a knowable position. A static shot of a person talking has neither. An arc around a running machine has both.
The moves that make it audible
Not all camera movement produces audible spatial change. The eight movement primitives that cover most of what people prompt split cleanly.
| Camera move | Spatial payoff |
|---|---|
| Orbit / arc around a subject | High — the whole soundscape rotates, this is the clearest demonstration |
| Pan or whip pan across a scene | High — sources sweep across the field as they leave and enter frame |
| Tracking alongside a moving subject | Medium — the world moves past, the subject stays anchored |
| Dolly in / push toward a source | Medium — near-field sound gets closer, the bed recedes behind it |
| Dolly out / pull back | Medium — the reverse; useful for reveals |
| Tilt up or down | Low — vertical position carries much less than horizontal |
| Crane / boom | Low to medium, depending on what is under the camera |
| Locked-off | None — nothing moves, so nothing pans |
The pattern: horizontal movement past a fixed source is what you hear. Vertical and forward movement change distance and perspective, which is subtler and easier to miss on small speakers. If you are spending prompt effort on this at all, spend it on an orbit or a pan.
One thing to plan around: models have a strong prior toward adding movement, so a shot you intended as locked-off will often drift unless you say so. "The camera is entirely motionless" is the phrasing that gets a static frame — omitting movement language does not.
Writing the prompt
The shot skeleton that works generally works here too: cinematography first, then subject, action, context, and style. Putting the shot size and move at the front changes what you get; putting them at the end often gets them dropped. The VEO 3.1 prompting notes go through the full structure.
For spatial audio specifically, three additions:
Name the source and put it somewhere. "Espresso machine hissing at frame left" gives the model a position to anchor. "Café sounds" gives it a bed with nothing to move. The named, placed source is the entire mechanism — no anchor, no pan.
Name one camera move, unambiguously. Contradictory pairs like "static handheld" or "locked-off tracking shot" make the model pick one at random, and if it picks the static one you lose the effect entirely. One move, stated once. Note also that speed and easing are never reliably inferable from text: "slow orbit" is re-estimated on every generation, so if the sequence needs the same move at the same speed across several shots, camera control as a parameter beats camera control as an adjective.
Keep the audio instructions in their own sentences. One idea per sentence — the move in the cinematography line, the source in an ambience or SFX line. Stacked into one clause they get parsed as a single blurred request.
Putting it together:
Slow 180° orbit around a woman seated at a counter, medium shot, 35mm. She looks past the camera and does not speak. Ambient noise: an espresso machine hissing at frame left, low chatter throughout the room. SFX: a cup sets down on the counter beside her. No music.
The orbit is the move, the espresso machine is the anchor, the cup is the one event that will read regardless of playback device.
When it earns the extra work
Spatial audio is worth prompting for when the answer to "will anyone hear this" is yes. That is narrower than it sounds.
Worth it:
- Long-form and mid-form content watched with headphones, where the stereo field survives.
- Reveals — something audible before it is visible, then the camera finds it. The spatial move is doing narrative work, not decoration.
- Establishing shots at the top of a piece, where you are selling a sense of place.
- Orbits and arcs where the subject is the still point and the world moves. This is the shot the feature was built for.
Not worth it:
- Talking heads. One source, centred, static — the effect has nothing to act on.
- Product-on-white and other single-source shots.
- Anything that mostly autoplays muted. If the sound-off version carries the video, sound-on refinements are a bonus at best. Designing for sound-on versus sound-off is the decision that comes first.
- Fast-cut sequences. A pan needs a couple of seconds to register as movement; on a 1.5-second cut it reads as a jump.
Spatial audio is a finishing touch that costs prompt tokens and generation attempts, competing with the shot itself for the same budget. Where the picture already works, it adds polish. It never rescues a shot.
The phone-speaker check
Most short-form video is watched on a single small speaker, which means the stereo field is summed to mono before anyone hears it. A sound panned hard left survives that sum — it just arrives with no position. On a phone speaker, the pan is simply gone.
That is a reason to run the check rather than skip it, because two things can go wrong in the fold-down:
- Anything that relied on width thins out. A wide, phase-decorrelated bed can partially cancel when summed, and the mix loses body. The symptom is a bed that sounds full in headphones and hollow on a phone.
- The thing you actually needed becomes inaudible. If the sound carrying the moment was placed off to one side and the centre is empty, mono fold-down leaves you with a quiet, ambiguous middle.
The check itself takes a minute:
- Play the clip on a phone speaker at arm's length, in a room with normal background noise.
- Confirm the one sound that matters — the line, the impact, the reveal cue — is audible without the pan doing any work.
- Then listen on headphones to confirm the spatial move is there and worth having.
Do this on a preview rather than a paid render. The editor is EDL-based and preview: true gives a 480p pass at no credit cost with a short per-user cooldown between them; audio comes through the downscale unchanged, so a preview answers this completely. The final export is charged once regardless of clip count. If the phone test fails, the fix is not more spatial detail — it is a centred, dry version of the important sound, with the spatial material demoted to the bed.
FAQ
Which models support audio that moves with the camera?
None of them advertise it, so you test rather than assume. Start from models that generate audio at all: the catalog's audio-capable list filters to models whose own feature tags name an audio capability, and the native-audio model map sorts them into ambience, synced effects and dialogue. Then run one orbit around a single named source and listen in headphones. A couple of generations settles it.
Will the panning survive if I re-encode or add music in the edit?
The stereo image survives a normal export. What flattens it is a mono deliverable, and what buries it is a music bed set too loud. If you have gone to the trouble of generating spatial audio, keep the added bed well under it and check in headphones before committing to the mix.
Can I get a specific direction, like "the truck passes right to left"?
You can steer it by describing the source's position and the camera's move, and the model will usually honour the geometry. What you cannot rely on is exact timing or speed, because both are re-estimated every generation. Treat direction as reliable and timing as approximate, and put anything that must land on a precise frame into the edit as a separate effect.
Is spatial audio worth it for vertical short-form at all?
Occasionally, for a reveal or an establishing shot at the top. Mostly not, because single-speaker playback removes the effect and fast cutting removes the time it needs to register. The budget is better spent on making the one important sound clear and centred. What generating video with sound involves covers the cost side of that trade.