A mixed track is a sum, and un-summing it is genuinely hard: the parts overlap in frequency and time, so a model has to infer what each source was rather than filter it out. Modern separation is good enough that the resulting stems are usable in an edit, which was not true a few years ago.
For video work it opens up ordinary things that used to need the original session files. Drop the vocal and keep the instrumental as a bed under a voiceover. Keep the vocal and rebuild around it. Duck one element under dialogue instead of the whole track.
Artefacts appear where sources overlap most. Cymbals bleeding into a vocal stem, a bass note ghosting under the drums — usually inaudible in a full mix, sometimes obvious when a stem is soloed and turned up.
In practice
- Judge stems in the mix you will actually use, not soloed.
- Instrumental beds from separation are the common case for voiceover work.
- Rights do not change because a track was separated — the underlying licence still applies.
The mistake to avoid
Treating separation as a licensing workaround. Removing a vocal does not make someone else's recording yours to publish.
Where you will run into it
- Split a Song Into Stems — One track in. Two usable stems out.
- Isolate Vocals From a Track — Strip the instrumental. Keep the voice.
Related terms
Voice isolation
Voice isolation separates speech from everything else in a recording — traffic, room noise, music — and keeps only the voice.
Text-to-music
Text-to-music generates an original composition from a written description of genre, instrumentation, mood and tempo — with or without sung lyrics.
Speech-to-text
Speech-to-text converts spoken audio into written text, producing the transcript that captions, translation and search all depend on.
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
Voice cloning
Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.