When Spatial Audio Is the Wrong Deliverable
Spatial and binaural masters do not survive most social and web players. When a 5.1 bed is worth the fold-down, and when stereo is the honest master.
"Spatial" is three products sharing a brochure. Binaural is a headphone illusion. 5.1 is six channels with a layout and a fold-down. Stereo that pans with a camera is still two channels. Shipping the wrong one as the master does not make a social cut feel expensive. It makes a player fold your mix in a way you never heard, or it makes a binaural file sound hollow on a TV.
Music generators and native-audio video models will happily sell a sense of space. Treat space as a delivery format with a destination, not as a quality upgrade on the stereo bed you will actually post. If the destination is a feed, a site, or a phone speaker, the honest master is stereo, checked in mono. If the destination is a platform that ingests discrete 5.1 and you can monitor the fold-down, then a 5.1 bed can earn the work.
Three different products, one word
Keep the labels separate or you will QA the wrong file.
| Deliverable | What it is | What plays it correctly | What happens everywhere else |
|---|---|---|---|
| Stereo (Lo/Ro) | Two channels, left and right | Phones, laptops, the web, social apps, most TVs | This is everywhere else |
| Binaural | Two channels encoded with a head-related transfer function, for ears | Headphones, some spatial players that expect binaural | Speakers scramble the cues; it can sound phasey or thin |
| 5.1 | L, C, R, Ls, Rs, plus optional LFE | A 5.1 (or better) monitoring chain and a pipeline that keeps discrete channels | The player downmixes to stereo or mono with whatever coefficients it has |
ITU-R BS.775-3 is the reference for the ear-level 5.1 layout: fronts on a 60° arc, surrounds behind, LFE an optional band-limited enhancement rather than "the subwoofer." The same recommendation gives downmix coefficients so six channels can become two. A common Lo/Ro shape puts the centre into left and right at 0.7071 (-3 dB) so dialogue does not jump 6 dB when C folds. The LFE is not required to survive a stereo downmix; many folds discard it. If the weight of the track lived in LFE, the stereo you never listened to is anaemic.
Binaural is not a downmix of 5.1. It is a different encode. A binaural render that sounds like a room on headphones can collapse on a soundbar because those cues were meant for pinnae, not for two drivers in a bar. Do not deliver binaural as "the stereo master" and hope.
Object-based formats need a renderer, metadata, and a destination that accepts them. A WAV named spatial_final.wav in a social export preset is not that pipeline.
Versely's music path is a stereo generate. generate_music does not expose a 5.1 or binaural layout, and the music generator is the right tool for a stereo bed. If another model or a DAW offers a spatial or binaural bounce, treat it as a second master for a second destination, not as an upgrade of the file you will attach under a VO.
When 5.1 is worth the fold-down
A 5.1 bed earns its keep when three things are true at once.
The destination actually ingests discrete 5.1 (or better) and plays it that way for a real share of the audience. Broadcast, some OTT masters, a festival DCP, a linear TV cut: those are jobs with a spec. A YouTube upload that can carry 5.1 for some listeners is not automatically a 5.1 job. If you cannot name the ingest spec, you do not have one.
You can monitor it. Five speakers at the BS.775 angles, plus a decision about LFE, in a room that is not a laptop. Mixing 5.1 on headphones with a binaural fold is how centre dialogue gets written too quiet (because the headphone fold already put it in both ears) and how surrounds get written too loud (because they feel exciting in a simulated room). If the only chain you have is a stereo desk, write a stereo master.
You will listen to the downmix as a deliverable, not as a glance. Author the Lo/Ro (or the Lt/Rt, if that is what the spec names) and play it on the stereo chain you trust, then on a phone. Centre-heavy dialogue should still sit. Surrounds should not dump a second music bed onto the VO. LFE should not have been the only bass.
The work is not upmixing a stereo generate with a spatializer and calling it 5.1. An upmixer invents a centre and a surround from a record that was already wide. Real 5.1 for picture is stems: DX in the centre, MX as a bed you placed, FX attached on purpose. If you never cut those stems, you have a stereo job in a 5.1 container.
Native-audio video does not give you those stems. Native audio is a stereo print from the same pass as the frames. Models that generate sound with picture are the right pick for a one-language clip that should feel like a room. They are not a 5.1 bus.
When stereo is the honest master
Default to stereo for social, web, most ads, and anything that will be watched on a phone, a laptop, or a TV that the viewer did not calibrate.
- Feeds. TikTok, Reels, Shorts, and similar players are stereo-or-less in practice. A 5.1 file gets folded by someone else's coefficients. A binaural file gets played as if it were ordinary stereo.
- Web players and landing pages. Two channels, often through laptop speakers. Your added music should be a stereo bed that still exists in mono.
- Talking-head and VO-led spots. The voice is the programme. Space that the viewer cannot hear is mix budget you could have spent on intelligibility.
- Anything you cannot monitor in the format you would ship. Unmonitored 5.1 is a guess. A monitored stereo print is a decision.
Stereo is not a downgrade. It is the format that matches the device. Make it compatible: centred voice, bass in the mid, a bed that survives a ten-second mono fold, true peak under the ceiling you actually mean (EBU R 128 publishes -1 dBTP). That file travels. A "spatial" bounce that you only heard on headphones does not.
If a client asks for "spatial" because a deck used the word, ask which of the three products they mean, and which player will be the judge. Most of the time they mean the VO should not sit in a vacuum. That is ambience and a sparse stereo bed, which you can generate as added audio. It is not a surround master.
Native-audio pan is not a surround bed
A native-audio model that keeps an espresso machine on the left while the camera orbits is doing source-anchored stereo. When that pan is worth the prompt is a headphones-and-narrative question: orbits, reveals, establishing shots. It is still two channels.
Do not confuse that effect with 5.1. There is no centre, no surround pair, no LFE, and no catalog switch that declares a spatial specification. You find out by generating an orbit and listening. For a social master, fold to mono for ten seconds and confirm the machine still exists. For a 5.1 master you need stems, five speakers, and a downmix you signed off on.
The same caution applies when a music model offers a spatial or binaural extra render. Use it for the place that asked for it. Do not attach that extra as the bed under a 15-second ad.
House rule: one master per destination, named after the destination. spot_social_stereo.wav goes under the VO. spot_51.wav exists only if a spec asked for it and you heard the fold-down. Copying one file to three names is how the wrong product ships.
FAQ
Is a binaural bounce a good "universal" master?
No. Binaural is a headphone product. Speakers do not reconstruct the head it was encoded for. If you need one file that fails least often, that file is Lo/Ro stereo, mono-checked, not a binaural render.
Does ITU-R BS.775 mean I have to mix at those speaker angles at home?
If you are delivering 5.1, you need a monitoring layout that resembles the recommendation, or you are inventing a surround image nobody else will hear. If you are delivering stereo, you need a stereo seat (and a phone-speaker pass). Do not build a fake 5.1 in a plugin and skip the room.
Can I upmix a Versely stereo generate to 5.1 for a client who asked for surround?
You can, in a DAW, with an upmixer. What you get is a stereo record spread into more speakers. Listen to the stereo downmix of that upmix. If it is worse than the original stereo, you added a failure mode. The honest response, when you only have stereo generates, is a stereo master plus a real 5.1 mix if the spec is real.
Does Versely export 5.1 or binaural from the editor?
The editor composites a timeline and exports a video with a stereo soundtrack. Music and speech generates are stereo. Native-audio clips are stereo files that may image left-to-right. None of that is a 5.1 stem session. If a job needs discrete surround, the surround work happens in a DAW against a spec, after you have stems.