A peak meter answers one narrow question: did any sample get too loud. Loudness answers the one people actually mean when they say a clip sounds quiet — it is measured across the whole file, weighted the way human hearing actually responds to different frequencies, and expressed as a single LUFS number. Two files can share an identical peak and still read as very different loudnesses, which is why a peak-only mix so often comes out quiet next to someone else's.
The number is routinely treated as one shared target, and it is not. EBU R128 is a real, published standard — an integrated programme-loudness target of −23.0 LUFS, sitting in the open for over a decade — but it is a broadcast figure, not a universal one. Other destinations publish their own, different numbers, and at least one major platform does not publish an integrated target at all. Mixing every destination to a single borrowed figure, often one repeated in a forum rather than read from a spec, is how a technically-competent mix still fails a specific platform's own check.
For generated voiceover and music specifically, the number that matters is the finished mix, not the isolated generation. A cleanly-measured voiceover track can still land quiet or clipped once it sits under music and effects, because loudness is a property of everything playing together, not of any one layer in isolation.
In practice
- Measure the finished mix, not the isolated voiceover or music generation — loudness is a property of the whole track.
- Match the destination's own published figure rather than one number repeated across every platform — several genuinely differ, and at least one publishes none at all.
- A loud-sounding solo take and a loud-sounding mix are different claims; layering music and effects changes the integrated measurement.
The mistake to avoid
Normalizing to a single number heard secondhand in a forum or a template, rather than to the actual published figure for the destination you are delivering to. At least one major platform has never published an integrated target at all, which makes the commonly-repeated number for it a guess wearing a unit.
Go deeper
Loudness Normalization: Why Your Video Sounds Quiet
One organization publishes an exact, measurable loudness standard. Platforms apply their own and don't document it. Knowing which is which changes how you mix.
Where you will run into it
- Add a Voiceover to a Video — Type the script. Get a narrated video back.
- AI Music Generator — Describe a vibe. Get a song. Keep the rights.
Related terms
Voiceover
A voiceover is a spoken track laid under picture, rather than lipsync that drives a mouth or a caption track that visualises speech.
Text-to-music
Text-to-music generates an original composition from a written description of genre, instrumentation, mood and tempo — with or without sung lyrics.
Native audio
Native audio means a video model generates its own soundtrack — dialogue, effects, ambience — in the same pass as the picture, rather than leaving you a silent clip.
Stem separation
Stem separation splits a finished mix into its component parts — vocals, drums, bass, other instruments — as separate audio files.
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.