Free forever · Privacy-first

    Free Vocal Remover

    Drop a song, a voiceover with a music bed, or a video clip, and get the voice and the music back as separate tracks. HTDemucs, Meta's open music separation model, runs inside your browser: choose vocals + instrumental, or four stems (vocals, drums, bass, other). Solo and mute them in the built-in mixer, then download 16-bit WAV. Up to 6 minutes per file, nothing uploaded.

    • Looks paidActually free
    • No uploadRuns on your device
    • No accountNothing to sign up for
    • No watermarkYour file, untouched

    More free audio tools

    Same quality bar. Still free. Still no upload.

    Your files never leave your device

    Free to use, with no account and no watermark. The first run downloads the model files once and your browser caches them; the model then runs on your device, and the file you process is never sent anywhere.

    This is not a promise about how carefully we handle your upload. There is no upload. The file you choose is read into memory by your own browser, processed there, and handed back as a download — the tool has no server component and makes no network request with your data. The only download is the model itself, which contains nothing of yours.

    • Nothing is stored. We never receive the file, so there is no copy to retain, no retention period to disclose, and nothing to delete on request.
    • Nothing is transmitted. No upload endpoint, no analytics payload carrying file contents, no third-party processor.
    • You can verify it. Run the tool once so the model is cached, then disconnect from the internet and run it again. It still works, because everything it needs is already on your machine.
    • Nothing persists after you close the tab. Files are held in page memory only — not in local storage, not in a cache, not in a database.

    A note on wording, because it matters: we do not describe these tools as “encrypted.” Encryption protects data that travels to somebody else's computer. Your file does not travel, so there is nothing to encrypt and nothing to intercept — which is a stronger guarantee than encryption, not a weaker one. The page itself is served over HTTPS like the rest of the site.

    This makes the tools safe for material you are contractually barred from uploading to third-party services — client footage under NDA, unreleased campaign assets, anything covered by a confidentiality clause.

    How it works

    1. 1. Drop the file

      MP3, WAV, M4A, FLAC, or a video (MP4, MOV, WebM, up to 1 GB), up to 6 minutes. It is decoded in this tab; nothing is uploaded to Versely.

    2. 2. Pick two or four stems

      Vocals + instrumental for an acapella and a backing bed, or four stems for vocals, drums, bass and everything else.

    3. 3. Split on your device

      The model runs in the background with WebGPU, or on the processor with WebAssembly when WebGPU is missing. The first run downloads the 174 MB model once into your browser cache. Stop works at any time.

    4. 4. Mix and download

      Solo, mute and set the level of each stem, then download any stem, or the current mix, as a 44.1 kHz WAV.

    Questions

    Is this vocal remover really free?

    Yes: no account, no watermark, no per-minute charge. The separation model runs on your own device, so there is no server cost to pass on. The limit is 6 minutes per file, because the stems are held in your browser's memory while you mix them.

    Is my song uploaded anywhere?

    No. The file is decoded and separated inside this browser tab. The only network requests are the first-run download of the open model from Hugging Face (174 MB, then cached by your browser) and the ONNX Runtime engine from the jsDelivr CDN.

    Which model does it use, and can I use it commercially?

    HTDemucs (Hybrid Transformer Demucs v4) from Meta AI, released under the MIT licence, in an ONNX export of the official weights. The MIT licence covers the model. It gives you no rights over the music you put through it: stems of a commercial song are still that song, so clear it before you publish.

    How clean are the stems?

    Clear lead vocals over a band usually come out well. Expect faint bleed: reverb tails, backing vocals, and instruments in the same range as the voice (a lead synth, a sax) can leak into the vocal stem, and loud cymbals can leave traces in the instrumental. It is an estimate, not the original multitrack.

    How long does it take?

    With WebGPU (current Chrome, Edge and Safari) a 3-minute song took about 25 seconds in our tests on an Apple M4 Pro; older or integrated graphics take longer. Without WebGPU it runs on the processor: about 1.5 times the track length on that same machine, several times on older laptops. It runs in the background, so the page stays usable, and Stop cancels at once.

    Can I remove music from a video and keep the dialogue?

    Yes. Drop the video, choose vocals + instrumental, and the vocal stem carries the speech. Download it as WAV and lay it back under the picture in your editor. Sound effects and room noise land in the instrumental, so check the result before you replace the original track. Phones may run out of memory on long clips; a laptop is safer.

    Related tools

    Stay in the same job cluster. Free tools stay on-device; paid tools are labelled as such.

    These rearrange files. Versely makes new ones.

    Everything on this page works on a file you already have, which is why it costs nothing. Generating something that did not exist (video from a prompt, a voice, a score) runs a model, and that costs credits.