Your voiceover sounds muddy under AI music
AI music beds crowd the same 250–500Hz band as speech. Carve that range, set a level relationship, and apply them in the right order.
Turning the music down is not the first fix. If the voiceover and the bed are both dense between 250 and 500Hz, they occupy the same space at any level. You can hear every word and still feel like you are listening through a wall. That band is where speech has its body, and it is where a lot of generated tracks put warm pads, boxed-in midrange synths, and "lo-fi" mud on purpose. Two files, one fight.
The mix that works is a specific carve on the music, a level relationship that keeps the bed well under the voice, and an order of operations that does not bake the collision in before you can EQ it.
The voice and the bed are fighting for the same band
Spoken voice carries most of its weight in the low-mids. 250–500Hz is the body. 4–7kHz is the presence and the "AI sheen" on generated instruments. A text-to-music track that was prompted as "warm, full, cinematic" will cheerfully fill both. Suno-lineage outputs in particular tend to arrive already mastered, already loud, already occupying the middle of the spectrum. They were not arranged as a bed. They were arranged as the record.
That is why a static fader move feels like it should work and does not. You are lowering a brick that is still the same shape as the voice. Ducking the music under a voiceover covers the Versely-side volume controls (attach_audio_to_video in mix mode, music_volume versus original_volume). This article is the spectral half of the same problem: what to do when the fader is already low and the words still feel muddy.
A quick listen that names the fault:
| What you hear | Band | Fix |
|---|---|---|
| Voice is audible but "in a box," no chest, no air | 250–500Hz collision | Cut that range on the bed, not the voice |
| Harsh, fizzy top on the music, sibilance fights it | 4–7kHz | Gentle cut on the bed's presence, de-ess the voice separately |
| Words disappear only on kicks / bass notes | Masking on transients | Duck (sidechain) after the carve, or mute the drum stem |
| Bed has a ghost singer under your VO | Stem leakage | Use the instrumental as a bed, do not surgical-EQ a leaked vocal |
Do not carve the voice to make room for the music. The voice is the programme. The bed moves.
Carve 250–500Hz on the music, not the voice
Versely's attach-audio path is a static mix, not an EQ. The carve happens in a DAW or in your NLE's EQ on the music clip. Keep the voiceover flat until the bed is out of its way.
A working starting point, then listen:
- High-pass the bed around 80–100Hz if there is no musical bass you need. Rumble under a voiceover is extra mud.
- Cut 250–500Hz on the bed, a wide dip, about 3–6dB. This is the carve. You are not deleting the track. You are making a hole the voice already occupies.
- Tame 4–7kHz on the bed if the generate has that glassy mid-presence that generated instruments love. A dB or two. If you overdo this the bed goes dull and you will be tempted to turn it up, which puts the mud back.
- Leave the voice alone except for a high-pass around 80Hz and a de-ess if isolation hyped the sibilants.
Sweep the 250–500Hz cut while the voice is playing. The right frequency is the one where the words pop forward without the bed becoming a telephone. If you solo the bed, the carve will sound wrong. Judge it under the voice.
Generated tracks that still fight after a 6dB dip are usually too busy, not too loud. Prompt the next generate for space rather than EQ-ing a brick into a bed:
"Instrumental only, no vocals, sparse piano and light side-stick, leave a hole in the low-mids for a voiceover, no wide warm pad, no bass-heavy mix, 45 seconds."
The AI music generator will still try to make a finished record. "Sparse" and "instrumental only" are the words that change the spectrum. "Cinematic, lush, full" are the words that refill 250–500Hz.
If the track you already like has a sung hook, split it. separate_music_vocals on a Versely-generated Suno track returns vocal and instrumental stems. Use the instrumental as the bed. Stem separation is inferred, not recorded: drum stems carry ghost vocals, vocal stems carry instrument reverb tails. Use stems for mute-the-singer and balance, not for surgical EQ on a leaked syllable. Uploaded files are a different split (isolate_audio); that tool is not the Suno stem path.
Set a level relationship, then duck if you still need it
After the carve, set a relationship, not a vibe.
A working range: the music sits roughly 18 to 25dB under the voice. That is a lot quieter than it feels when you solo the bed. Soloing the bed is how muddy mixes get made. Listen to both, on a phone speaker and on headphones.
On Versely, attach_audio_to_video in mix mode keeps the original (or the voiceover) and layers the bed with separate music_volume and original_volume controls. Those are static for the whole clip. There is no sidechain in that call. For a 20–40 second talking head with no instrumental gaps, static is enough. Pull music_volume further than you think, then stop.
When the edit has holes (a pause, a cutaway, a hook with no VO), static will feel like the music vanished. That is when you sidechain in an NLE:
- Duck the bed 9–12dB under speech.
- Attack under about 300ms, so the first syllable is not buried.
- Slow release, so it breathes rather than pumping on every consonant.
Do the duck after the carve. Sidechaining a dense 250–500Hz brick just makes a pumping dense brick. Sidechaining a carved bed makes the voice sit.
Programme loudness is a separate pass, last:
- Streaming / social: about −14 LUFS, true peak −1 dBTP.
- Podcast: −16 to −19 LUFS.
Those numbers are for the finished mix, voice plus bed. If the mix is loud and the words still feel small, the bed is eating the loudness budget. Lower the bed.
Always master from WAV, never from the MP3 of the generate. Bounce to WAV, carve, level, duck, then encode the delivery file once.
The order: isolate, generate sparse, carve, level, master from WAV
Doing these out of order is how you "EQ" a problem that is still a prompt, or duck a problem that is still a spectrum.
- Isolate the voice if it is not a clean TTS bounce. Voice isolation first, so transcription, lipsync, and this mix are all looking at the same vocal. Do not isolate a vocal that was already clean; the model will start eating the body you need.
- Generate or pick the bed with space in it. Instrumental, sparse, no competing vocal. Extend to the video length rather than looping. Add music to a video is the attach step; extend a music track is the length step. Length before level.
- Carve 250–500Hz (and, if needed, 4–7kHz) on the bed in the DAW or NLE. Not in the prompt, not on the voice.
- Set the static relationship. Bed 18–25dB under voice. On Versely, mix mode,
music_volumelow. This is the add a voiceover plus music path, not replace mode, which would throw the voice out. - Duck only if the edit has gaps that should bloom. 9–12dB, fast attack, slow release.
- Master the WAV to the LUFS / true-peak target for the destination. Then encode.
A single agent request that respects the Versely-shaped part of this (generate, length, static mix) and leaves the carve where it belongs:
"Generate a sparse instrumental bed, no vocals, no warm pad, for this 40-second voiceover. If it's short, extend it to the video length. Mix it under the voice with music_volume well below the voice, mix mode not replace. I'll EQ the bed myself. If the generate still has a sung hook, split stems and attach the instrumental only."
The editor's 480p preview pass is free, with a short per-user cooldown, and the final export is a single charge. Preview will tell you if the bed is still sitting on the words. It will not show you a 250Hz build-up as a picture. Wear headphones for this QC. Phone speakers hide mud and then the upload sounds like a blanket.
FAQ
Why did turning the music down make it worse?
You lowered a dense midrange brick and then turned it back up because the bed "disappeared." The collision is the shape, not the fader. Carve 250–500Hz first. A quiet, full bed is mud. A quieter, carved bed is a bed.
Can I fix this by isolating the voice harder?
If the voice was recorded under the music, isolation can recover it. If the voice is already a clean TTS file, more isolation will thin the 250–500Hz body you need and the bed will win. Isolation is for noise around a voice, not for a mix you have not carved.
Should I use the vocal stem of the AI track under my voiceover?
No. That is two voices. Use the instrumental. If the instrumental still has a ghost of the sung line, that is stem leakage: keep the level low, do not EQ the ghost as if it were a real singer, and generate a new instrumental-only bed if it is obvious.
Is Versely's mix a sidechain?
No. attach_audio_to_video in mix mode is a static music_volume / original_volume for the whole clip. That is the right tool for a short talking head with continuous VO. For ducking that moves with speech, carve and sidechain in an NLE after you attach, or accept a static bed that stays low.