You need the prior generate's identifier, not an arbitrary MP3. A continue-at time picks where the new section should pick up; prompt and style steer the added bars. The point is to keep the take you already liked instead of rolling a new song and hoping the mood matches.
Video-extend continues picture from the last frame — a different medium, same idea of continuation. Text-to-music is the original write. Duration on a video model is a menu of clip lengths, not a way to stretch audio. Time-stretching a file you already have is an edit, not this job, and it will chip the transients.
Keep the new prompt close to the original settings if you want a seam instead of a gear change. The editing surface is extend a music track; the agent job is extend a song.
In practice
- Look up the original generate and extend that record; an upload is the wrong input shape.
- Pick a continue point mid-phrase rather than on a hard fade, so the model has motion to continue.
- If the bed is right and only short, extend it; if the bed is wrong, write a new one.
The mistake to avoid
Feeding an uploaded MP3 into extend, or rolling a new song because the first fade was early. Extend needs the generated record; a new prompt is a different track.
Go deeper
Extending a song continues a generated track you already have
extend_music needs a prior Suno audioId. A new generate_music prompt is a different job, and using it as an extend wastes credits.
Where you will run into it
- Extend a Music Track — Not long enough? Keep it going.
- AI Music Generator — Describe a vibe. Get a song. Keep the rights.
Related terms
Video extend
Video extend continues an existing clip past its final frame, generating new footage that starts from where the old footage stopped.
Text-to-music
Text-to-music generates an original composition from a written description of genre, instrumentation, mood and tempo — with or without sung lyrics.
Duration
Duration is how long a generated clip runs, chosen before generation from whatever lengths the model supports rather than trimmed afterwards.
Cover song
A cover song, in generation, is a new arrangement of an existing track you point at, not an original composition from a blank prompt.
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.