Guides

    Fold Down Surround Beds Without Losing Center

    Generated music parked in 5.1 L/R starves dialogue on stereo TVs. Fold down with center protected and LFE discarded for social and web masters.

    Versely Team7 min read

    A 5.1 bed that parks generated music in Left and Right will starve the dialogue on a stereo TV. The fold-down does exactly what the coefficients say. Center is attenuated. LFE is either thrown away (correct) or summed into small speakers (wrong). Neither outcome is mysterious once you look at the matrix.

    Most generated music is a stereo file. Generate a song from a prompt and you get a two-channel bed, not a broadcast surround stem. The failure starts when that stereo file is dropped onto a 5.1 timeline: the NLE puts energy in L and R, leaves Center empty, and maybe derives a rumble into LFE. Dialogue, mixed properly, lives in Center. A stereo television then plays the fold-down, not your surround room. Music stays at unity in L/R. Dialogue arrives through the Center coefficient, quieter than you mixed it. The TV did not "lose the voice." You asked the matrix to shrink it.

    What ITU actually sums

    Recommendation ITU-R BS.775 is the named surround layout for broadcasting, with five main channels (L, R, C, LS, RS) and an optional LFE. Annex 4 gives the stereo (2/0) downmix of those five main channels:

    Output L R C LS RS
    L′ 1.0000 0.0000 0.7071 0.7071 0.0000
    R′ 0.0000 1.0000 0.7071 0.0000 0.7071

    0.7071 is −3.01 dB. LFE is not in that table. The recommendation is explicit that the LFE channel is optional at the receiver and should carry only extra enhancement, and that LFE is typically discarded in a two-channel downmix so it does not overload small stereo speakers.

    Read the Center column against a real mix. If dialogue is only in C, it arrives in the stereo pair 3 dB down. If generated music is only in L and R, it arrives at 0 dB. The stereo TV then has a music-loud, voice-quiet mix even though the 5.1 room sounded balanced. Surrounds (LS/RS) also enter at −3 dB. If an upmixer parked more music in the surrounds, that music still survives the fold-down. Center is the channel the matrix treats as optional. Dialogue is not optional.

    ITU-R BS.1770 loudness measurement also excludes LFE. A hot LFE does not move your LUFS number and still blows up a laptop speaker if you include it in the fold. Discard it for social and web. Loudness normalisation is the programme-loudness document; it does not rescue a fold-down that already buried the voice.

    Why generated music lands in L/R

    A stereo file on a 5.1 bus has to go somewhere. Default routing is L/R. Center stays silent unless you put something there. That is correct for a music-only surround mix in a cinema. It is wrong for a talking-head brand film that will be watched on a stereo television, a phone, or a browser.

    Native audio video models do not publish a 5.1 specification. Do not treat a spatially interesting stereo render as surround. Spatial audio that pans with the camera is source-anchored stereo behaviour, and it is inaudible on a phone speaker. Folding a fake 5.1 of that render will not make the espresso machine orbit. It will make the voice recede.

    Upmixers that promise "5.1 from stereo" often do the same dump: L/R get the original, Center gets a summed remnant or nothing, LFE gets a low-passed copy. That is a width trick. It is not a dialogue-safe bed.

    The mix to build instead, on the 5.1 bus, if you truly need a surround master:

    • Center: dialogue, dry enough to read, no music except a deliberate, low-level mono sum if you want the fold-down to keep voice-to-music ratio.
    • L/R: music and stereo effects, already ducked under the voice you will hear in Center.
    • LS/RS: ambience, not the melody. If the melody lives in the surrounds, the fold-down will still sing over the words.
    • LFE: optional effects only, never the only place a rumble exists. If the rumble matters on a TV, it has to live in the main channels too.

    Ducking music under a voiceover still applies. A static music_volume under original_volume in Versely's mix is the short-form version of the same idea. Surround does not exempt you from it. The AI music generator and add music to a video are how the bed gets onto the picture in stereo. Surround, if required, is a later bus, not a generation setting.

    A fold-down recipe that protects Center

    Do not send 5.1 to a social platform or a web player and hope. Mix a stereo master. Name it as a deliverable.

    For social and web

    1. Create a stereo bus. Do not use the receiver's default matrix as your master.
    2. Copy Center into L and R at 0 dB, not at −3 dB. Dialogue is the programme.
    3. Add L and R music/effects at the levels they already had after ducking.
    4. Add LS and RS at −6 dB (0.5) or omit them. If you hear the song in the surrounds, omit them. ITU's own menu of surround coefficients includes 0.7071, 0.5000 and 0.0000 for a reason.
    5. LFE at 0. Discard it. If you miss the rumble, put that rumble in L/R of the stereo master as a filtered copy, then true-peak the mix.
    6. True-peak ceiling −1 dBTP. Measure integrated loudness against the standard you named on the spec sheet, on the stereo result, not on the 5.1 bus.

    If you must fold an existing 5.1 file

    Replace the Center coefficient. ITU's 0.7071 is the compatibility default, not a law that dialogue has to obey. A finishing fold for web:

    L′ = L + 1.0·C + 0.5·LS + 0·LFE
    R′ = R + 1.0·C + 0.5·RS + 0·LFE
    

    Listen. If the voice still recedes, the music in L/R is too loud, which is a duck, not a coefficient. Pull the bed, or split the generated track into stems and rebuild; split a song into stems is the escape hatch when the bed and a vocal in the music file are glued together.

    What you should hear on a phone

    Play the stereo master on a phone speaker and on a TV in stereo mode. If the words recede, the fold failed, whatever the 5.1 room said. If the rumble vanished and you needed it, it was only in LFE and you were always going to lose it. Put it in the mains.

    Keep the 5.1 mix if a client named a surround deliverable. Ship the stereo fold as a separate file with its own line on the spec sheet. One file cannot be both masters unless you have heard the fold on a stereo TV and signed it.

    FAQ

    Should I generate the music in 5.1?

    Not unless a tool you can name actually writes six discrete channels and you have heard them. Versely's music generation is a song from a prompt, attached as audio. Treat the result as stereo, duck it, and only upmix if a surround master was in the brief.

    What about the LFE rumble I wanted?

    If it only exists in LFE, stereo TVs and phones will not play it. Put a filtered copy in the main channels of the stereo master, then check true peak. Discard LFE on the fold either way. BS.775's own warning is overload on small speakers.

    Do I apply the ITU −3 dB Center coefficient for web?

    Not if Center is where the words are and L/R is where the generated music is. That coefficient is how dialogue loses to the bed. Use 0 dB on Center for the stereo master, then duck L/R until the words are comfortable.

    Can I leave fold-down to YouTube or the smart TV?

    You can, and you will not control the matrix. Some devices discard LFE, some do not. Some attenuate Center, some collapse everything. A stereo master you mixed is the only fold you can QC. Send that for social and web.