Speech, voice and audio

    Voice isolation

    Also called Audio isolation, Noise removal, Vocal extraction.

    Voice isolation separates speech from everything else in a recording — traffic, room noise, music — and keeps only the voice.

    It is separation rather than filtering. Older noise reduction attenuated frequency bands and took some of the voice with them, which is the source of that thin, underwater quality on over-processed audio. A model trained to separate sources reconstructs the voice as its own signal instead, so the noise can go entirely while the speech stays full.

    The wins are largest on material you cannot re-record: an interview beside a road, a phone recording in a busy room, archive footage. It is also a preprocessing step — transcription, dubbing and lipsync all work better on an isolated vocal than on a full mix.

    There are limits worth knowing. Reverb is part of the voice's own signal and is harder to remove than noise sitting behind it. And aggressive isolation on already-clean audio makes things worse, because the model starts removing parts of the speech it is unsure about.

    In practice

    • Run it before transcription or dubbing, not after.
    • Reverb and clipping resist isolation; steady background noise does not.
    • Compare against the original at full volume — over-processing is easiest to hear on breaths.

    The mistake to avoid

    Isolating audio that was already clean. There is nothing to remove, so the model removes detail from the voice instead.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.