Describe the vibe: generate_music returns a track, not a music video
Generate a song from a prompt is one agent job. generate_music gives you audio. Lyrics-only, SFX, and picture are different jobs.
Describe the vibe. Get a track. Generate a song from a prompt runs generate_music and stops at audio. A lo-fi instrumental, a pop vocal about a city, a moody score for a product reveal — those are this job. Asking it for a music video, a whoosh, or lyrics you have not approved yet spends a music generation on the wrong artifact.
The tool is generate_music. Required: a prompt. Optional: instrumental, title, style, vocal gender, duration, and the usual Suno steering. You get a finished track you can attach later, extend, cover, or split into stems. The AI music generator is the same generate with a form.
Audio is the product
generate_music does not open a timeline. It does not animate a still to the downbeat. It does not write a riser that lands on a cut. If you needed a hit, a footstep, or rain on glass, that is generate a sound effect (generate_sound_effect). If you needed the words first so you can edit the chorus before you commit to a vocal take, that is write song lyrics (generate_lyrics) — text only, no audio, on purpose.
The waste pattern is a prompt like "make a cinematic whoosh song for the title" into music, then another music call when the result is a bed with a vocal you did not want, then an SFX call you should have made first. Each of those is a generation. The capability prices per generation, shown before you confirm.
Use this page when the brief is a track: genre, mood, instrumental or not, vocal gender if it sings. "Upbeat lo-fi hip hop instrumental for a study playlist" is a clean ask. "Write and generate a pop song about chasing your dreams, female vocals" is a clean ask. "Score this 30-second ad and cut the video to it" is at least two jobs: music, then attach.
Attach is later
Once you have a keeper, add music or a voiceover to a video lays it. replace if the original sound should die. mix if you are bedding under existing voice. Generating a second full song because the first was two seconds too short is the expensive way to extend; the music tools include an extend path for a prior Suno audioId.
Do not ask generate_music to invent picture. A music video from a song is a different production, not a longer prompt.
FAQ
Will generate_music also make the video?
No. You get audio. Picture is a video job. Pairing them is attach or a timeline render, after both files exist.
When should I write lyrics first?
When you care about the words. generate_lyrics is text. Review the chorus, then generate the song. Committing to a vocal take you will rewrite is a wasted music generation.
Is a whoosh a song?
No. Short SFX is generate_sound_effect. Music models will try to be musical. That is the wrong texture for a UI click or a title riser.
Can I reuse one track across a series?
Yes, if the series should share a bed. Generating a new custom song per episode is how a channel loses a recognizable sound and how the music meter runs for no brand reason.