ElevenLabs Music stems for an edit
When you need to duck the bed under VO, stems beat a stereo bounce. Music v2 is the stem job.
A stereo bounce is a finished record. An edit needs parts. If the voiceover has to sit on the track, "turn it down" is not a mix — it is a fight with a file that was mastered as the song. Stems are how you stop that fight before it starts.
ElevenLabs Music v2 and Dubbing v2 is the June 2026 decode: chunk-based composition plans, performance-conditioned dubbing, the v1 speech sunset. This post does not retell that. It is one production rule. When the brief is a bed under a voice, Music v2 is useful because you can get stems, not because the bounce sounds more "cinematic."
Why a stereo bounce loses the edit
A generated song arrives loud, full, and already wearing a vocal. Put a founder VO on top and you have two leads and one loudness budget. Pulling the fader down leaves the sung hook poking through consonants. EQ-ing a hole in a bounce carves drums, pad and leftover vocal together. You are mixing a poster.
Stems split the poster into layers you can mute.
ElevenLabs documents stem separation on an audio file: a ZIP of separate stems, with variations that include a two-stem split (vocals and instrumental), a four-stem split (vocals, drums, bass, other), and a six-stem split for deeper control. Music v2 is the generator; the stem endpoint is the editor's door. The point of v2 for this job is not the rap delivery demo. It is that you can take a licensed, section-planned track and pull the lead out instead of prompting "instrumental" and hoping.
force_instrumental on compose is the other door: generate with no vocal in the first place. Use it when you know there will be VO. Use stem separation when you already like a vocal track and need the bed from it, or when you need drums and bass on separate faders.
The edit you actually do
Three cases, one rule: never duck a lead you cannot mute.
1. Talking-head, continuous VO. You need an instrumental bed. Generate Music v2 with instrumental forced, or split a song you already have and keep the instrumental stem. Then set a static relationship — bed well under voice — and stop. Dynamic ducking is optional on a clip with no gaps.
2. VO with holes. Cold open, pause, end card. You want the bed to bloom in the gaps. That is ducking, and ducking a stem is what the next post is for. Ducking a stereo bounce with a sung hook in it is how the hook reappears every time the voice stops.
3. You need the drums quieter than the pad. Four- and six-stem splits exist for this. If you only need "no singer," two stems are enough. Do not pay for a six-stem split to solve a vocal problem.
Do the split before you attach the file to picture. Attaching the bounce, hating the mix, then splitting is two exports and a timeline that now points at the wrong clip.
What Versely already does, and what it does not
Versely does not host ElevenLabs Music. Original tracks in the catalog run through the AI music generator (Suno Sounds V5.5). For a Suno track generated in-app, split a song into stems (separate_music_vocals) returns vocal and instrumental from that track's taskId and audioId. That is the two-stem job for Versely-native music. It is not a general splitter for an arbitrary MP3; uploaded files go through isolation instead.
So the practical map is:
| Source | Stem path |
|---|---|
| ElevenLabs Music v2 | ElevenLabs stem separation (2 / 4 / 6) |
Versely generate_music |
separate_music_vocals (vocal / instrumental) |
| A random upload | Isolation, not Music-v2 stems |
If the client bought ElevenLabs for the licence and the stem export, do the music there, bring the instrumental (or the drum-less bed) into the edit, and do not bounce a full mix "for simplicity." Simplicity is how the singer ends up under the VO.
Stem separation is inferred, not a studio session. Drum stems carry ghost vocals. Vocal stems carry reverb tails of the band. Use stems to mute a part and to balance, not to surgical-EQ a leaked syllable and call it a master.
When you should not stem
- The track is the content. A music video, a chorus-led Short, anything where the vocal is the hook: you want the bounce. Stems are for when speech is the programme.
- You can generate instrumental from the start. Forcing instrumental is cheaper than generating a song and splitting it. Stem the keep; do not stem as a default.
- You need a true M&E for broadcast. A two-stem AI split is not a recorded music-and-effects deliverable. If the SOW says M&E, budget a real mix, not a separator.
Music v2's composition plans still help this job: intro / build / drop / outro as chunks means the bed can be quiet under the VO and move when the VO stops, without a loop seam. That is structure, which the June post covers. The only addition here is: structure is wasted if the file you cut is a stereo brick.
FAQ
Is a Music v2 instrumental generate the same as a stem split?
No. force_instrumental never writes a vocal. A stem split removes a vocal that was already there, and the instrumental will still carry ghosts of it. If you know there is VO, generate instrumental. If you fell in love with a vocal version, split it and listen for leakage before you duck.
Can I stem a Versely Suno track the ElevenLabs way?
No. Versely's splitter only sees in-app generate_music / extend_music IDs. ElevenLabs' stem endpoint sees a file you send it. Do not mix the two pipelines on one cue. Pick the generator, then use that generator's split.
Do I need six stems for a 30-second ad?
Almost never. Two stems solve the singer-under-VO problem. Four stems if the kick is masking speech and you want the pad to stay. Six stems are an arrangement job, not a UGC ad job.
Does stemming change the licence?
Stem separation is a transform of a track you already generated. It does not mint a new licence. If the original generate was commercially cleared on that product, the stem is still that generate. If you are mixing Firefly beds with ElevenLabs vocals, that is a different terms question — read which model wrote which layer, the same way you would for speech.