Speech, voice and audio

    Stem separation

    Also called Source separation, Stems.

    Stem separation splits a finished mix into its component parts — vocals, drums, bass, other instruments — as separate audio files.

    A mixed track is a sum, and un-summing it is genuinely hard: the parts overlap in frequency and time, so a model has to infer what each source was rather than filter it out. Modern separation is good enough that the resulting stems are usable in an edit, which was not true a few years ago.

    For video work it opens up ordinary things that used to need the original session files. Drop the vocal and keep the instrumental as a bed under a voiceover. Keep the vocal and rebuild around it. Duck one element under dialogue instead of the whole track.

    Artefacts appear where sources overlap most. Cymbals bleeding into a vocal stem, a bass note ghosting under the drums — usually inaudible in a full mix, sometimes obvious when a stem is soloed and turned up.

    In practice

    • Judge stems in the mix you will actually use, not soloed.
    • Instrumental beds from separation are the common case for voiceover work.
    • Rights do not change because a track was separated — the underlying licence still applies.

    The mistake to avoid

    Treating separation as a licensing workaround. Removing a vocal does not make someone else's recording yours to publish.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.