Guides

    Loudness Normalization: Why Your Video Sounds Quiet

    One organization publishes an exact, measurable loudness standard. Platforms apply their own and don't document it. Knowing which is which changes how you mix.

    Versely Team7 min read

    "Why does my video sound quiet next to everyone else's" has two possible answers and creators usually reach for the wrong one first. The first answer is a real, published, numeric standard that's been sitting in the open for over a decade. The second is a number that gets repeated constantly in creator forums, sounds exactly as authoritative as the first, and was never actually published by the platform it's supposedly targeting. Mixing against the wrong one of those two is a genuinely common way to end up quiet without ever touching your fader.

    Two different bodies answering two different questions

    The European Broadcasting Union publishes a specific, testable loudness target that any broadcaster or creator can read and implement exactly. Streaming platforms apply their own loudness normalization on ingest — pulling hot uploads down toward some internal target — without publishing the algorithm or the exact number in their own documentation. One of these is a standard. The other is a black box that's been reverse-engineered by creators running test uploads and comparing before-and-after levels, and the two get treated as equally authoritative online constantly, which is where most of the confusion starts.

    What R128 actually specifies

    EBU R 128 sets Programme Loudness at a Target Level of −23.0 LUFS — and it's worth reading the standard's own two tolerances rather than the single rounded number that usually gets quoted. Where hitting the target exactly isn't practically achievable — live programming is the standard's own example — a tolerance of ±1.0 LU is permitted, with the explicit caution that a broadcaster shouldn't let deviation toward the edge of that range become standard practice. Separately, for implementing loudness workflows in quality-control environments, a tighter ±0.2 LU tolerance is allowed specifically to account for measurement error between meters. Those are two different numbers for two different problems — one for the reality of live production, one for the precision limits of the meters checking the work — and conflating them is how "the target is roughly −23" turns into an argument about whose meter is wrong.

    The standard also sets a True Peak Level ceiling of −1 dBTP, measured against a meter compliant with ITU-R BS.1770 and EBU Tech 3341, with its own ±0.3 dB measurement tolerance — this is the number that actually prevents clipping regardless of what any platform's normalizer does to the file afterward. And the loudness measurement itself runs a relative gate at −10 LU: passages quieter than 10 LU below the ungated average are excluded from the calculation, so a long silent or near-silent stretch doesn't drag the reported loudness down and make a genuinely loud programme measure as quiet on paper.

    None of these numbers are R128 inventing its own physics. ITU-R BS.1770 is the underlying measurement algorithm — R128 cites it directly as the standard any compliant meter has to implement — and it's what actually defines LUFS and LU as units in the first place. R128 is the target and the tolerances built on top of that measurement machinery, not a competing way of measuring loudness.

    The short-form supplement, and why it exists separately

    A fifteen-second ad and a fifty-minute broadcast aren't the same measurement problem, and the EBU addresses that directly rather than expecting one document to cover both. R 128 s1 — titled, plainly, "Loudness Parameters for Short-Form Content (Adverts; Promos, etc.)" — is the supplement R128 itself points to for exactly this case, and the reason a separate document is warranted is duration: some of the standard's own measures, Loudness Range chief among them, are explicitly not recommended for anything under about a minute, because there simply isn't enough data in a fifteen-second clip for that particular measure to mean anything. A promo mixed to the letter of the full R128 document but ignoring its short-form supplement is following most of the spec and missing the part written specifically for content its own length.

    Where the number everyone actually quotes comes from

    Ask around and you'll hear "−14 LUFS" offered as gospel for YouTube, Spotify and half a dozen other platforms in the same breath as R128's −23. It's worth being direct about the difference: that figure is not published in YouTube's own Help documentation. It's a community-measured estimate, arrived at by creators uploading test tones and controlled mixes and observing what the platform's own normalization did to them on playback — useful as a working approximation, and genuinely not the same category of fact as a number written into a formal recommendation with a document history and a drafting committee behind it. Treat platform loudness targets as best-effort field measurements that can drift as platforms quietly retune their own normalization, not as a spec you can cite the way you'd cite R128.

    Why your video sounds quiet even though you never touched the fader

    Most platform loudness normalization only really works in one direction. A hot mix gets pulled down toward whatever internal target the platform is normalizing to — that's straightforward gain reduction, with no real risk attached. A quiet mix getting boosted back up is a different, riskier operation, because pushing gain up on something that's already using most of its headroom risks running past true peak and clipping — so platforms tend to normalize down aggressively and normalize up conservatively, if at all. The practical result: a video mixed conservatively, or mixed to broadcast-safe −23 LUFS and then dropped into a platform silently expecting something closer to community-measured −14, doesn't get rescued by that platform's own normalizer. It just sits quieter than everything mixed hotter around it, exactly matching the original complaint.

    A defensible target when the platform's own number isn't published

    Without a published figure to mix toward for most social platforms, the more defensible move is mixing to R128 s1's short-form parameters as your documented baseline and treating community-measured platform figures as a rough sanity check rather than the actual target — and keeping true peak under that −1 dBTP ceiling regardless of which number you mixed the integrated loudness toward, since that's the one figure in this whole picture that's both published and protects you no matter which platform's normalizer touches the file next.

    A Versely walkthrough

    Loudness is a mixing decision made before audio ever reaches Versely's editor — the tool that matters once a track's ready is attach_audio_to_video, in whichever mode the fix actually calls for. If the complaint is genuinely "too quiet" on a track that's otherwise balanced, replacing the video's audio outright with a properly leveled version is the direct fix. More often, though, "sounds quiet" turns out to mean "the dialogue is quiet relative to the music," which isn't a loudness problem at all — it's a balance problem, and stem separation is the actual tool for it: split the existing mix into vocal and instrumental stems, rebuild with the instrumental pulled back under the dialogue, and the perceived loudness issue often disappears without the overall level moving at all.

    For adding a new element on purpose, add music to a video or add a voiceover using attach_audio_to_video in mix mode, setting the original track's volume down rather than leaving two full-level sources fighting each other:

    "Mix this voiceover onto the video. Duck the existing background music under it — I want the voice clearly up front, not competing with the track."

    That's the practical version of the loudness discipline this whole piece is about: get the balance and the levels right going in, mix toward a documented target rather than a forum number when one's available, and keep true peak under control — because no platform's normalizer is going to rescue a mix that was quiet, or unbalanced, before it ever left the editor.