Your ducking pumps instead of breathing
Pumping is a release-time problem, not proof sidechain was wrong. Attack, release and depth settings that hide the duck behind the voice.
The duck is working. That is why the mix sounds worse. Every time the voice lands, the music slams down and then slams back up in the gap, so the bed gasps in time with the sentences. Operators often treat that pumping as proof that sidechain was the wrong call, and they flatten the music to a static whisper that never returns. The compressor is doing the job you asked. The release is just too fast for speech.
Static level is a different problem, and ducking music under a voiceover covers that case: two files playing at once, neither moving. Pumping only appears once something is actually riding the bed against the voice. If you hear the music bounce, you are in a compressor, an auto-duck, or a keyframed fade that recovers too quickly. Fix the envelope. Do not throw the sidechain out.
What pumping actually is
A sidechain compressor on the music listens to the voice. When the voice crosses the threshold, gain reduction pulls the bed down. When the voice stops, the compressor lets go and the bed comes back. Three numbers decide whether that motion disappears behind the words or becomes a second performance:
- Attack is how quickly the bed yields once speech starts.
- Release is how quickly the bed returns once speech stops.
- Depth (range, or the combination of threshold and ratio) is how far the bed drops.
Pumping is a release-time artefact. A fast release is the EDM trick: the kick punches a hole and the rest of the mix blooms back between hits, on purpose, in time with the tempo. Speech is not a kick. Words have irregular gaps. A release that recovers in those gaps produces a gulp after every phrase. Slow the return so the swell feels like a breath in the pause, not a bounce on every consonant.
Attack errors sound different. If the attack is slow, the first syllable of each line sits under full-level music, then the bed finally drops. That is masking, not pumping. Depth errors are obvious suction: the whole track disappears, then floods back. Get the three controls pointed at the right failures before you touch EQ.
The three controls that hide the duck
Use this as a starting patch on a compressor or auto-duck that has a dedicated range/depth control, then listen. These are mix starting points, not a broadcast spec.
| Control | Start here | If you overshoot |
|---|---|---|
| Attack | Fast enough that the first syllable is clear, and in any case under about 300 ms | Slow attack buries the start of each line |
| Release | Clearly slower than the attack: several hundred milliseconds, so the bed swells through the gap | Fast release pumps between words |
| Depth | 9–12 dB of gain reduction on speech | Much more than that reads as the track being muted |
| Floor | Music sitting roughly 18–25 dB under the voice while someone is talking | If you still notice the music during speech, it is still too loud |
The 9–12 dB depth and the under-300 ms attack are the numbers that actually matter for spoken-word ducking. The 18–25 dB under-voice figure is the static relationship while speech is present, not the gap level. In the pauses the bed should come up. That swell is the whole point of a moving duck.
A few extra compressor habits stop chatter:
- If the plugin has a hold, use a short hold so a quick breath does not release the duck and immediately grab again.
- High-pass the sidechain input around the body of the voice so a kick, a chair scrape, or a bass note in the music does not trigger the duck. You want speech, not every transient in the session.
- Do not also compress the music heavily for loudness on the same insert. The sidechain needs headroom to move. Loudness is a later, slower decision. Why your video sounds quiet is the place for programme targets, not the duck.
A starting patch, then two listen tests
Set it once, then stop staring at the gain-reduction meter. The meter will always look busy on speech. Your ears have to decide whether the motion is audible.
Test 1: the first syllable. Loop a sentence that starts with a plosive or a stressed word ("Buy this", "Stop scrolling", a product name). If the music is still at full level on that first sound, shorten the attack. Stay under about 300 ms. You should not hear a fade; you should hear the voice arrive in a hole that is already there.
Test 2: the gap after a sentence. Loop a line that ends, then sits for half a second before the next line. The bed should rise through that gap, not snap. If it snaps, lengthen the release. If it never rises, your release is so long that the next sentence arrives before the bed has moved, which is a static duck wearing a compressor’s clothing. Shorten it until the pause has a swell and the next line still ducks in time.
Then check depth. 9 dB is usually enough for a sparse bed. Dense generated tracks, especially ones with their own vocal hook, often need the top of that 9–12 dB range, or they need a different bed. If you are past 12 dB and the voice is still fighting, the compressor is compensating for a frequency collision, not a level collision. Turn the depth back to 10 dB and go to EQ.
A one-line version of the patch, for a plugin that exposes attack, release, and range:
Attack 20 ms, release 500 ms, range 10 dB, threshold just under the voice peaks, ratio 4:1. High-pass the key input. Listen on a phone speaker. Lengthen release until the gulp disappears.
Twenty milliseconds and five hundred milliseconds are a place to start, not a law. The constraints that do not move: attack under about 300 ms, release slower than attack, depth in the 9–12 dB band.
When the problem is frequency, not the compressor
Generated beds are dense in the same 250–500 Hz region as a speaking voice. A compressor ducks level. It does not carve a pocket. If the kick, the pad, and the voice are all sitting in that band, a perfect envelope still leaves a muddy fight, and you will keep reaching for more depth until the mix pumps.
Do the carve on the music, not the voice:
- A wide cut in the 250–500 Hz band on the bed, a few dB, is usually enough to let the voice’s body through.
- If the track has a sung vocal, do not duck that vocal into being an instrumental. Split it. Stem separation on a generated track, via splitting a song into stems, gives you an instrumental bed you can actually duck. Treat those stems as separated from a mixture, not as isolated mics: drum stems often carry ghost vocals. Use them for mutes and broad level, not surgical EQ.
If the bed is shorter than the video and you looped it, the loop point will also read as a pump. Extend the music track before you touch the compressor again.
Do this in a DAW, then bring the mix back
Versely’s in-app mix is a static relative level. attach_audio_to_video in mix mode exposes music_volume and original_volume for the whole clip, not a sidechain. That is the right tool when speech is continuous and you only need the bed quieter than the voice. It will not produce pumping, because nothing is moving. It also will not swell in the gaps.
When you actually want a moving duck, mix it in a DAW or an NLE that has a compressor with a sidechain input. Export a single WAV of the mixed soundtrack, then replace the video’s audio with that file, or add music only after the bed is already riding correctly.
The order that avoids rework:
- Generate or pick the bed. Keep it WAV, not an MP3 bounce.
- Fix length (extend, do not loop).
- Carve 250–500 Hz if the voice and the bed share that space.
- Set the moving duck: attack under ~300 ms, slow release, 9–12 dB depth, music 18–25 dB under speech.
- Print the mix, attach it, stop touching the compressor.
The AI music generator gets you the bed. The duck is a mix decision you make before you attach it. If the result still gulps, the release is still too fast.
FAQ
Is pumping a sign I should stop sidechaining?
No. Pumping means the bed is returning too quickly after each phrase. Lengthen the release until the rise sits in the pause and disappears under the next line. A static quieter bed hides the gulp and also hides the music. Keep the sidechain; fix the envelope.
How far should the music drop under the voice?
While someone is talking, a 9–12 dB duck is the working depth, with the bed sitting roughly 18–25 dB under the voice. In the gaps it should come back up. If you need more than about 12 dB of gain reduction for the words to be comfortable, the collision is probably in 250–500 Hz, not in overall level.
Why does Versely’s mix not pump?
Because it is not a sidechain. Mix mode applies a fixed music_volume against a fixed original_volume for the duration of the clip. Useful for short talking-head clips with almost no pause. If you want the bed to breathe in the gaps, print a ducked WAV in a tool that has attack and release, then replace the video audio with that mix.
Should I duck the whole track or just the mids?
Duck the whole bed for speech, then carve 250–500 Hz on the music so the compressor is not doing frequency work. Multiband sidechain on only the voice band is a valid second step on a dense generated track, but it is a refinement. Get a slow-release, 9–12 dB wideband duck working first.