Guides

    Vocal isolation pulls a voice from any mix; it is not a Suno stem split

    isolate_audio takes any mixed audio URL. separate_music_vocals needs Suno ids. Using the wrong splitter wastes credits on a file it cannot split.

    Versely Team4 min read

    Mixed audio in. Clean vocals out. That is the job. Isolate vocals from an audio clip is one named agent capability. isolate_audio takes any audio_url — a field recording, a downloaded mix, a voice-over sitting on a bed — and returns a vocal-only track. It is not the Suno stem tool. Separate vocals from a song is separate_music_vocals, and that one needs taskId and audioId from a track you generated here. Pointing isolation at a Suno record you still have ids for, or pointing the stem tool at a Zoom recording, is how you spend a split on a file the tool cannot split.

    Stem separation and voice isolation are related ideas. On Versely they are two rows.

    What isolate_audio actually returns

    Give the agent the mixed clip. It isolates the vocal track and strips background music and instrumentation. You get back a dry voice file. Cost is priced per generation, shown before you confirm.

    Required field: audio_url. That is the whole surface. No taskId. No "also give me the instrumental." If you needed both stems of a generated song, you wanted the other capability. If you needed a transcript, you wanted transcribe audio to text — isolation does not type.

    The finishing door is isolate vocals from a track. This page is the agent verb. Isolated vocals are a new asset. Attach them, transcribe them, or drop them on a timeline later. They are not "the song, fixed."

    What this job is not

    A Suno vocal/instrumental split. separate_music_vocals is for a prior generate_music / extend_music result. It wants the ids the generator minted. Isolation on that same track still works as a mix-in, mix-out pass, but you throw away the instrumental stem the Suno splitter would have given you. If the track is a Versely song with ids, use the stem job.

    An extend. Extend a song makes the track longer. Isolation does not add bars. People isolate because the bed is loud, then generate a new song "without music" and pay twice. Duck the bed, or isolate, or regenerate as instrumental — pick one named job.

    A voice change. Change the voice in an audio clip re-voices a performance. Isolation does not swap the throat. Isolate first if the bed is in the way of a change-voice pass; do not ask isolation to be Cartesia.

    Video. This tool takes audio. A video file with a loud bed wants the audio pulled or a video-side attach/replace job, not a hope that audio_url will "see" the picture.

    Split the file you have, with the tool that matches it

    The waste is identity mismatch. A client WAV from a shoot is isolate_audio. A Suno bed you made at 2am is separate_music_vocals. Running the Suno tool on the client WAV fails the ids. Running isolation on the Suno track and then generating a replacement instrumental is a second song that will not lock to the vocal you just pulled.

    If the job is "I need the words with no music," isolation. If the job is "I need karaoke of the track we generated," stems. If the job is "I need a transcript," do not isolate as a prelude to transcribe unless the bed is actually drowning the speech model — transcribe is its own cheap-looking pass and a different row.

    Do not isolate every take "in case." Dry vocals of a miss are still a miss, now with a generation charge on top.

    FAQ

    Can isolate_audio split a Suno song I made in Versely?

    It can pull vocals from the mix because it takes any audio URL. You will not get the matching instrumental stem. If you still have taskId and audioId, separate vocals from a song is the named job that returns both.

    Does isolation give me an instrumental?

    No. Vocals out. The bed is what gets removed, not exported. Dual stems are the Suno splitter.

    Should I isolate before I transcribe?

    Only if the bed is drowning the speech. Transcription is transcribe audio to text. Isolation plus transcribe is two bills when one transcript would have done.

    Is this the same as removing music from a video?

    This capability is audio-in, audio-out. For a video, pull or replace the audio on the video jobs, or extract the audio first. Do not point isolate_audio at a video URL and expect a silent picture with dry voice already married back. That marriage is a different attach step.