Lyrics First: Reviewing Words Before You Render Music
Generated lyric text costs nothing to edit before you commit. Skip that step and every fix to one bad line means burning a full song generation to redo it all.
Asking directly for a finished song is a bet placed on two things at once: whether the words land, and whether the track sounds right — and a single generation gives you exactly one roll on both together. If the melody's great and the second verse says something you'd never actually put in a brand video, there's no way to keep the music and swap the line. The only lever is regenerating the whole thing and hoping the next roll gets both halves right simultaneously. Splitting the words out as their own step first doesn't change what the model can do — it changes what a bad result costs you, from a wasted generation to a five-second edit.
The bundled default is a single roll on two things at once
Asking for a song directly from a theme runs lyrics and composition through one generation, which means a lyric problem and a music problem are indistinguishable until you're listening to a finished track. Maybe the words are exactly right and the arrangement feels flat. Maybe the arrangement is great and one line reads wrong the moment you hear it sung out loud instead of imagined in your head. Either way, there's no partial fix available — the words and the music arrived fused, and the only tool for changing one is regenerating both.
What splitting the render actually buys you
generate_lyrics produces standalone lyric text from a theme or prompt and nothing else — explicitly no audio, built for exactly the moment before you'd otherwise commit a full song generation to words you haven't actually read yet. That's the entire value of the split: text is free to iterate on in a way audio never is. Read it, cut a verse that's dragging, swap a phrase that doesn't scan, run it again from a tweaked prompt if the whole angle missed — every pass costs a text generation, not a music one, and none of it commits you to a finished track until you're actually satisfied with the words on the page. Only once the lyrics read the way you want do they get pasted into generate_music to become the actual audio, which is the point where a bet on composition finally gets placed on words you've already approved rather than words you're seeing for the first time alongside the melody.
What to actually check before you commit
Reading lyrics as text rather than hearing them sung changes what you're able to catch, and it's worth using that difference deliberately rather than just skimming for typos. A few things read completely differently on the page than they will sung against a beat: whether a line's natural syllable count and stress pattern actually fits the genre you're about to generate into — a line that scans fine as prose can be genuinely awkward once it has to land on specific beats. Whether the rhyme scheme (or deliberate absence of one) is consistent enough to feel intentional rather than accidental. Whether the actual message — the thing a brand or campaign track is supposed to say — survives the metaphor it's wrapped in, since a lyric can sound good and still miss the brief entirely. And, worth checking specifically for anything commercial: whether any phrase reads differently than intended once it's the literal line a listener will hear repeated on a hook. None of this is available to catch after the fact once it's baked into a finished vocal take — it's only checkable while it's still plain, editable text.
Structure is still the part good lyrics don't fix
Approved lyrics solve the words problem, not the arrangement problem, and it's worth being honest that a second real limitation sits right behind the first one. Versely's own reference on text-to-music generation names this directly: "Structure is the weak point. Models produce convincing texture and a plausible groove more easily than a piece that develops — an intro that earns a chorus, a drop that lands where your edit needs it." Great, fully-approved lyrics can still land inside a generation that doesn't build the way a finished song should, because that's a composition problem, not a words problem, and reviewing the lyrics first doesn't touch it. The practical answer is the same one that reference gives for arrangement generally: generate longer than the piece needs and cut to the edit, rather than trying to prompt the model into hitting an exact structural beat on command.
Extending is its own read-before-you-commit moment
The same logic worth applying before the first generation is worth applying again before stretching a track that already exists. extend_music continues a previously generated song from — or near — a point in the source track, and it requires that source track's audioId, pulled from the earlier generate_music result, which means it's only usable against a song you've already generated and kept, not a fresh idea. Before firing it, the same question the lyrics-first step answers on the front end is worth asking again on the back end: what does this extension actually need to add — a second verse continuing the story, an instrumental bridge, a full second chorus — rather than treating "make it longer" as a request with only one obvious shape. An extension is still a generation, with the same one-shot-per-attempt economics as the original; thinking through what the added section is supposed to do before requesting it is cheaper than finding out after.
A Versely walkthrough: words, then music, then more if you need it
The concrete sequence starts with the theme, not the track: a prompt like "Write lyrics for an upbeat pop song about a summer road trip with friends" calls generate_lyrics and returns text only — no credits spent on audio yet. Read it, edit lines directly or ask for a targeted rewrite of just the second verse, and repeat that pass as many times as the words actually need. Once they're right, "Use these lyrics to generate the track" pastes the approved text into generate_music, which is the point where composition finally gets committed against words that have already been reviewed rather than words arriving for the first time inside a finished mix — from there, writing custom lyrics for background music or asking the agent to write song lyrics directly are the two framings of the same first step. If the finished track later needs to run longer — a video cut grew, or the song just needs a second chorus — extending the song is the follow-up move, referencing the original track's audioId and continuing from the point you specify rather than starting over.
Takeaway
A song generated directly from a theme bets on lyrics and composition at the same time, with no way to isolate which one actually went wrong. Pulling the words out as their own step first doesn't make the model better at either job — it just moves every lyric fix from "regenerate the whole track and hope" to "edit a line of text," which is a difference in cost, not in capability. Read the words while they're still free to change. Commit the generation once they're actually right.