Guides

    Mix a generated bed under the cut, not under candidate takes

    generate_music plus attach_audio_to_video in mix mode scores a file. Scoring a miss, or scoring before picture lock, is a bed you will mute.

    Versely Team3 min read

    Background music sits under whatever is already playing. generate_music writes a track from a prompt. attach_audio_to_video in mix mode layers it, with separate music_volume and original_volume. That pair is a score for a cut you are keeping. Mix it onto a candidate and you have paid for a bed, and maybe a song, that will never ship.

    Add music to a video is the job. The AI music generator is the writer. Attachment is a flat credit fee regardless of duration. The song is billed separately if you do not already have one.

    Mix is not replace

    replace strips the original soundtrack. That is a different task. If the clip has dialogue, native room, or a voiceover you want, mix is the setting. People grab replace because it "sounds cleaner" on a scratch take with ugly buzz. They have now destroyed evidence of whether the take had usable sound, and they have glued a new song to a picture they may still regen.

    Duration is the other trap. A 12-second keeper and a 30-second maybe want different beds. Generating the song first, then hunting for a picture that fits it, inverts the pipeline. Picture lock tells you the runtime. Runtime tells you whether you need extend_music at all.

    Do not generate a library of beds "so we have options" against a bin of unlocked takes. You will attach the wrong energy to the wrong file, then generate again when the keeper is shorter.

    Score after the cut exists

    1. The take survives review. Native audio is either the mix or clearly going away.
    2. If native stays, mix. If native goes, that is replace, still after lock.
    3. Write the bed to the locked duration and mood. Attach once. Duck with the two volume knobs, not a fantasy of a mastering chain.

    estimate_cost before you generate a long prompt. Attachment is the cheap half. The song is the half people run too early because it feels like production. Production on a miss is waste.

    If the first pass fades out before the outro, extend the Suno track from a known audioId. Extending a song you attached to a discarded take is a sequel nobody asked for.

    What "under" actually means

    The bed should lose to the voice. If you cannot hear the line, you mixed like a trailer. Short-form almost always wants low music_volume. You learn that on the keeper, in the real loudness of this VO, not on a silent generate you have not narrated yet.

    The AI video generator makes the picture. Music is not how you hide a bad one.

    FAQ

    Can I generate the song while the picture is still sampling?

    You can write a reference bed. Do not attach it. Attachment glues a file to a file. When the picture changes, you either live with a runtime mismatch or you attach again.

    Why is mix the default for this task?

    Because "add music" in production language means a bed under existing sound. Replace is for when the original track is condemned.

    Does attach_audio_to_video scale with length?

    It charges a flat fee. The leak is not a long file. The leak is doing it twice, plus a second generate_music when the mood no longer fits the new take.

    Should I split stems before attaching?

    Only after this track is the track. Stem splits need the in-app taskId and audioId. Splitting a discarded generate is two extra calls on a song you will not mix.