It is separation rather than filtering. Older noise reduction attenuated frequency bands and took some of the voice with them, which is the source of that thin, underwater quality on over-processed audio. A model trained to separate sources reconstructs the voice as its own signal instead, so the noise can go entirely while the speech stays full.
The wins are largest on material you cannot re-record: an interview beside a road, a phone recording in a busy room, archive footage. It is also a preprocessing step — transcription, dubbing and lipsync all work better on an isolated vocal than on a full mix.
There are limits worth knowing. Reverb is part of the voice's own signal and is harder to remove than noise sitting behind it. And aggressive isolation on already-clean audio makes things worse, because the model starts removing parts of the speech it is unsure about.
In practice
- Run it before transcription or dubbing, not after.
- Reverb and clipping resist isolation; steady background noise does not.
- Compare against the original at full volume — over-processing is easiest to hear on breaths.
The mistake to avoid
Isolating audio that was already clean. There is nothing to remove, so the model removes detail from the voice instead.
Where you will run into it
- Isolate Vocals From a Track — Strip the instrumental. Keep the voice.
Related terms
Speech-to-text
Speech-to-text converts spoken audio into written text, producing the transcript that captions, translation and search all depend on.
Stem separation
Stem separation splits a finished mix into its component parts — vocals, drums, bass, other instruments — as separate audio files.
Speech-to-speech
Speech-to-speech takes a recording of one person talking and re-renders it in a different voice, keeping the original performance intact.
AI dubbing
AI dubbing replaces a video's spoken audio with another language, usually keeping the original speaker's voice and optionally re-syncing their mouth.
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.