Video plus audio, combined losslessly: attach_audio_to_video does not write the track
Add music or a voiceover to a video is one agent job. attach_audio_to_video lays an existing file. Generating speech or music is a different job.
Add music or a voiceover to a video is one agent job. attach_audio_to_video lays an existing file. Generating speech or music is a different job.
Generate a song from a prompt is one agent job. generate_music gives you audio. Lyrics-only, SFX, and picture are different jobs.
A full mix under VO is a fight. Get stems, or generate the bed without a lead, then duck.
cover_music reinterprets a track you already have, then you attach it. Covering a maybe, or covering before picture lock, spends a generate on a bed the cut will not keep.
generate_music plus attach_audio_to_video in mix mode scores a file. Scoring a miss, or scoring before picture lock, is a bed you will mute.
extend_music continues a prior Suno track from an audioId. Extending before picture lock, or extending a song you will not mix, pays for seconds the cut will not use.
generate_lyrics returns text. generate_music then spends a track. Writing lyrics before picture lock, or scoring a miss with them, pays for words and a song the cut will not hold.
Mid-side generated music can cancel in mono on TVs and some phones. Run a ten-second fold, read a correlation meter, and collapse width without killing the bed.
A practical map of every AI model inside Versely — what each one is best at, when to pick it, and how to route prompts across models for the best result.