Edit with music and captions in one pass
One edit_video call can stitch, mix a bed, and burn captions. Preview 480p. Do not chain three tools for one ad.
One edit_video call can stitch, mix a bed, and burn captions. Preview 480p. Do not chain three tools for one ad.
Generate silent, score with music after, when dialogue is not the point. Native audio is optional, not mandatory.
Music that fights the voice kills retention. Duck hard under narration and stop romanticizing a hot mix.
Add music or a voiceover to a video is one agent job. attach_audio_to_video lays an existing file. Generating speech or music is a different job.
Generate a song from a prompt is one agent job. generate_music gives you audio. Lyrics-only, SFX, and picture are different jobs.
A full mix under VO is a fight. Get stems, or generate the bed without a lead, then duck.
cover_music reinterprets a track you already have, then you attach it. Covering a maybe, or covering before picture lock, spends a generate on a bed the cut will not keep.
generate_music plus attach_audio_to_video in mix mode scores a file. Scoring a miss, or scoring before picture lock, is a bed you will mute.
extend_music continues a prior Suno track from an audioId. Extending before picture lock, or extending a song you will not mix, pays for seconds the cut will not use.
generate_lyrics returns text. generate_music then spends a track. Writing lyrics before picture lock, or scoring a miss with them, pays for words and a song the cut will not hold.
Mid-side generated music can cancel in mono on TVs and some phones. Run a ten-second fold, read a correlation meter, and collapse width without killing the bed.
A practical map of every AI model inside Versely — what each one is best at, when to pick it, and how to route prompts across models for the best result.