What you say to the agent
No special syntax — just describe it like you would to a person.
What it does, step by step
- 1
Reference a track generated earlier with the song-generation capability (the agent looks up its taskId/audioId from your history).
- 2
It splits it into vocal and instrumental stems.
- 3
You get back both stems separately.
What it needs from you
- •The source track's task ID (from a prior Suno generation)
- •The source track's audio ID (from a prior Suno generation)
What comes back
Two audio files: the isolated vocal stem and the isolated instrumental stem.
What it costs
Priced per generation, shown before you confirm.
Under the hood
This is what the agent actually calls when you ask for it — real tools from its live surface, not marketing copy.
separate_music_vocalsSplit a previously generated Suno track into vocal/instrumental stems. Requires the source track's taskId and audioId (from a prior generate_music / extend_music result).
See it done in a real workflow
Morning Matcha Routine
UGC-style wellness reel — a clean-girl creator walks through her actual morning matcha ritual, then cuts to a cozy animated hero shot of the finished iced latte.
TrendingNYC Street Interview
Photorealistic street-vlog interview — Riley works three different NYC corners at golden hour asking strangers one question: "What's the wildest thing you've ever done?" Three candid OTS/two-shot clips with locked character references, real handheld energy and native spoken dialogue. Vertical 9:16.
TrendingPrimo Protein vs Other Brand
Pixar-style 3D comparison ad — the confident PRIMO PROTEIN pouch faces off against the tired old rival brand's tub across 11 talking clips: pasture vs dusty pantry, herb garden vs toxic lab, clean American lab vs grimy factory, chocolate-milkshake CTA vs "wet sand." ~40s vertical reel with native character voices.
Or start from a one-tap template
Frequently asked questions
What do I actually say to the agent to separate vocals from a generated song?+
Just describe it in plain English — for example: "Split that last song you made into vocal and instrumental stems" The agent handles picking the right tool and model from there.
What does the agent need from me first?+
At minimum: The source track's task ID (from a prior Suno generation); The source track's audio ID (from a prior Suno generation). Anything else it needs, it asks for before running.
What do I get back?+
Two audio files: the isolated vocal stem and the isolated instrumental stem.
Does this cost credits?+
Priced per generation, shown before you confirm.
You can also just ask for
Ask your Versely agent to separate vocals from a generated song
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.