Workflows

    Scoring a video from one still image

    Lyria 3.5 takes an image as a musical brief. A pipeline for scoring a cut from its key frame, and how to keep one cue consistent across a multi-clip edit.

    Versely Team9 min read

    Lyria 3.5 will take an image as the prompt. Google's model page puts it plainly: upload an image and ask Lyria to turn it into a track (Lyria on Google DeepMind). It got almost none of the attention that went to the vocal and tempo improvements in the 29 July 2026 Flow Music release, and it is the one that changes a workflow rather than a spec.

    The reason is the direction of travel. A normal scoring job runs text to music to picture: you describe a mood in words, get a track, then try to make it agree with footage it has never seen. Image-to-music lets you run picture to music to picture. The brief is a frame that is already in your edit, so the track starts from the thing it has to match instead of from your description of that thing.

    A frame carries the half of the brief that words are bad at

    Write down what you actually mean by "warm, nostalgic, slightly melancholy." Now hand that sentence to three different people and look at what comes back. The vocabulary of mood is low-resolution, and it is low-resolution in exactly the places that matter for matching music to picture.

    A still carries those places directly:

    • Colour temperature and grade. Not "warm," the specific warm you graded.
    • Contrast and density. A crushed-black night interior and a lifted, hazy exterior want different amounts of low end.
    • Subject and scale. One face in close-up implies an intimate arrangement. A wide landscape implies space and reverb.
    • Era and place cues. Wardrobe, architecture, signage, typography inside the frame.
    • Implied motion. A frozen mid-stride shot reads faster than a static composition, and that reads through into tempo.

    What a frame cannot carry is equally important, because this is where people get burned. A still says nothing about length, tempo, structure, whether there are vocals, or how the piece ends. Those are not aesthetic questions, they are edit questions, and the frame has no opinion on them. So the frame is half a brief. The other half is a short written constraint layer, which is exactly what tempo and duration control are for.

    Treat the still the same way you would treat a reference image in a visual generation: it fixes the look and leaves the mechanics to you.

    The pipeline, step by step

    1. Lock the cut length before you generate anything. You cannot ask for a duration you have not decided. If the edit is still moving, lock a working length and accept that you will regenerate once. A rough assembly with real timings beats a finished track that fits nothing.

    2. Choose the key frame, and it is not frame one. The frame you want is the one carrying the emotional claim of the piece, which is usually two-thirds of the way in. For a product piece it is the hero shot. For a narrative it is the frame at the turn. For travel or brand film it is the widest establishing shot with the best light. Opening frames are almost always the worst choice because they are deliberately neutral.

    3. Extract it clean. Ask the agent to extract frames from a video and pull the frame at full resolution, before captions, overlays, letterboxing or a logo bug are burned in. This matters more than it sounds. A frame with a bright caption bar across the lower third biases the brief toward whatever that graphic looks like, and you will get a track that is subtly reacting to your subtitle style.

    4. Write the constraint layer. Attach the frame, then add the mechanics the frame cannot express:

    [key frame attached]
    Score this image. 41 seconds exactly. 100 BPM.
    Instrumental, no vocals.
    One build starting around 0:18, resolve at the end — not a fade.
    Keep the low end out of the way of a male voiceover.
    

    Five lines. Length, tempo, vocal decision, structure, mix constraint. The image handles everything else, and anything you add about mood is usually redundant with the picture you just attached.

    1. Generate at the cut length, not a round number. Duration control exists so you stop trimming 60-second tracks down to 41 and calling the truncation an ending.

    2. Audition against the picture, never in a player. Music that sounds good on its own and wrong under footage is the standard failure, and you cannot hear it in a waveform. Drop the track onto the timeline and render a 480p check pass. In the Versely video editor, preview: true gives you that pass free with a short per-user cooldown, and the final export is charged once regardless of how many clips are on the timeline. The full breakdown of that split is in editor previews and final export.

    3. Duck, then export. Music level against voiceover is a separate craft problem with a real method behind it, covered in ducking music under a voiceover. Do it after the track is chosen, not as a way to rescue a track that does not fit.

    Holding one cue together across a multi-clip edit

    Here is where the workflow goes wrong for most people. You have eight clips, image-to-music works well, so you score each clip from its own frame. Now you have eight unrelated cues in eight different keys with eight different tempos, and the edit sounds like a compilation.

    The fix is a discipline, not a feature:

    One brief for the whole edit. Pick one frame for the piece. Not one per scene, not one per act. Every additional generation is a new key and a new arrangement logic, and there is no reliable way to make two independent generations agree.

    Generate long, then cut. If the piece genuinely needs a different feeling in the middle section, generate one track long enough to contain both and cut sections out of it. Everything you take from a single generation shares a key and a tempo by construction.

    Extend rather than regenerate. When the edit grows past the track, extend the existing track instead of starting again. The method and its seams are covered in extending music into seamless background tracks.

    Use stems for variation. Stem separation is the cheapest way to get three cues out of one generation: full arrangement for the build, pad and bass only for the quiet section, drums alone under the montage. Same key, same tempo, same instrumentation, obviously related, because they are the same recording.

    Freeze the constraint text. If you must run a second generation, change only the length. Every word you change in the constraint layer is a chance for the model to reinterpret the brief.

    Once the cue is settled, the assembly is ordinary timeline work. The agent can build a video from clips, music and captions in one pass, and because the editor is EDL-based, swapping the track later re-renders the same timeline rather than making you rebuild it.

    What to check before the final export

    Check The failure it catches
    Track length equals timeline length, to the frame A one-second overhang that fades under black
    Music events line up with picture cuts A build that peaks two seconds after the reveal
    Level under the loudest voiceover line, not the average Narration disappearing in one sentence out of twelve
    Section joins, heard at speed Seams that are invisible on a waveform
    Ending is resolved, not truncated The most common tell of a generated score
    Generation record kept: model, prompt, date, account The clearance question, asked six months later

    That last row is not an editing check, but it belongs on the same list because it costs nothing at export time and costs a lot to reconstruct afterwards.

    If what you need is a bed rather than a scored cue, the shorter path is the AI music generator in Versely, where a looping bed with tempo and key control is a couple of credits and no frame extraction is involved.

    FAQ

    Which frame should I use if the video has no obvious hero shot?

    Use the frame you would pick as the thumbnail. That instinct is already selecting for "the frame that represents the piece," which is the same job the musical brief needs. If the thumbnail choice is genuinely arbitrary, the piece probably has no emotional centre, and the music will not supply one.

    Can I attach more than one image?

    Treat it as one image per generation. Two frames pull the brief in two directions and the result tends to average them into something that matches neither. If two sections of the edit genuinely need different treatments, generate one long track from the stronger frame and use stem variants for the contrast.

    Does image-to-music work for a bed, or only for a scored cue?

    It works for both, but it is worth less for a bed. The whole point of a bed is that it does not react to the picture, so a brief derived from a specific frame is effort spent on a property you are about to throw away. Beds are better served by a text prompt with a tempo and a key.

    Should the image be the graded frame or the raw one?

    The graded frame, every time. The grade is a large part of what you are asking the model to hear, and a flat log frame describes a different piece of video than the one you are shipping.