Mix a generated bed under the cut, not under candidate takes
generate_music plus attach_audio_to_video in mix mode scores a file. Scoring a miss, or scoring before picture lock, is a bed you will mute.
Background music sits under whatever is already playing. generate_music writes a track from a prompt. attach_audio_to_video in mix mode layers it, with separate music_volume and original_volume. That pair is a score for a cut you are keeping. Mix it onto a candidate and you have paid for a bed, and maybe a song, that will never ship.
Add music to a video is the job. The AI music generator is the writer. Attachment is a flat credit fee regardless of duration. The song is billed separately if you do not already have one.
Mix is not replace
replace strips the original soundtrack. That is a different task. If the clip has dialogue, native room, or a voiceover you want, mix is the setting. People grab replace because it "sounds cleaner" on a scratch take with ugly buzz. They have now destroyed evidence of whether the take had usable sound, and they have glued a new song to a picture they may still regen.
Duration is the other trap. A 12-second keeper and a 30-second maybe want different beds. Generating the song first, then hunting for a picture that fits it, inverts the pipeline. Picture lock tells you the runtime. Runtime tells you whether you need extend_music at all.
Do not generate a library of beds "so we have options" against a bin of unlocked takes. You will attach the wrong energy to the wrong file, then generate again when the keeper is shorter.
Score after the cut exists
- The take survives review. Native audio is either the mix or clearly going away.
- If native stays, mix. If native goes, that is replace, still after lock.
- Write the bed to the locked duration and mood. Attach once. Duck with the two volume knobs, not a fantasy of a mastering chain.
estimate_cost before you generate a long prompt. Attachment is the cheap half. The song is the half people run too early because it feels like production. Production on a miss is waste.
If the first pass fades out before the outro, extend the Suno track from a known audioId. Extending a song you attached to a discarded take is a sequel nobody asked for.
What "under" actually means
The bed should lose to the voice. If you cannot hear the line, you mixed like a trailer. Short-form almost always wants low music_volume. You learn that on the keeper, in the real loudness of this VO, not on a silent generate you have not narrated yet.
The AI video generator makes the picture. Music is not how you hide a bad one.
FAQ
Can I generate the song while the picture is still sampling?
You can write a reference bed. Do not attach it. Attachment glues a file to a file. When the picture changes, you either live with a runtime mismatch or you attach again.
Why is mix the default for this task?
Because "add music" in production language means a bed under existing sound. Replace is for when the original track is condemned.
Does attach_audio_to_video scale with length?
It charges a flat fee. The leak is not a long file. The leak is doing it twice, plus a second generate_music when the mood no longer fits the new take.
Should I split stems before attaching?
Only after this track is the track. Stem splits need the in-app taskId and audioId. Splitting a discarded generate is two extra calls on a song you will not mix.