Guides

    Video plus audio, combined losslessly: attach_audio_to_video does not write the track

    Add music or a voiceover to a video is one agent job. attach_audio_to_video lays an existing file. Generating speech or music is a different job.

    Versely Team3 min read

    Video plus audio, combined losslessly. Add music or a voiceover to a video is the lay: attach_audio_to_video takes a video_url and an audio_url you already have. It does not write a voiceover. It does not compose a song. Using it as a generate, or generating a new clip because you wanted different sound, spends either a missing-file retry or a video-model bill on a soundtrack problem.

    The tool has two modes. replace (default) strips original sound and substitutes yours — "mute the original," "put this music on the video." mix keeps the original and layers the new track, with music_volume and original_volume. Mix is "bed under the talking clip." The video stream is copied, not re-encoded. Fast. Flat credit fee regardless of length. The editing pages add music to video and add voiceover to video are the same split.

    Files first

    If you do not have the track, make the track. Speech is write and generate a voiceover (generate_speech). A bed is generate a song from a prompt. SFX is its own generate. Then attach. Asking attach to "add a warm narrator" with no audio URL is not a hidden TTS switch; it is the wrong job.

    Picture should already be the picture. A new take has a new duration. Replace on a recut will overhang or die early just like a mismatched VO. Mix on a take whose room is noise will not rescue the take. Decide replace vs mix on the keeper, not in the abstract.

    Do not reach for a full edit_video timeline when the cut is already right and you only need a soundtrack swap. Attach copies pictures losslessly. A timeline re-render re-encodes. That is the point of this job.

    Volume is a parameter

    Mix without setting volumes is how a bed eats the VO or vanishes. Say quieter. Say VO stays. Regenerating the music model to "make it quieter in the composition" is a generation for a fader.

    FAQ

    Does attach_audio_to_video generate the voiceover?

    No. Pass audio_url. Generate speech or music in their jobs, then lay. One attach fee, not a speech bill pretending to be attach.

    Mix or replace when the clip has native-model audio I dislike?

    Replace if that audio should die. Mix if you are adding a bed under words you want to keep. Native-audio garbage is a replace, not a louder bed.

    Will this recut the video to the song?

    No. Optional trim_to is not a beat-matched edit. Length mismatch is your problem to trim, on picture or on audio, as its own step.

    Why not regenerate the video with generate_audio on?

    Because you already have pictures you like. A new generate is a new plate. Attach keeps the plate and changes the ear.