AI Models

    Lyria 3.5: when a full song beats a looped bed

    Lyria 3.5 shipped into Flow Music on 29 July 2026 with direct tempo and duration control. When a three-minute vocal earns its place, and when a bed wins.

    Versely Team9 min read

    Google shipped Lyria 3.5 into Flow Music on 29 July 2026. Most of the coverage led on vocal quality. The item that actually changes how you work is further down the list: you can now state the tempo and the duration you want, on a model that generates up to three minutes with vocals.

    That sounds like a settings change. In practice it decides which shape of track you reach for. Most video work has quietly settled on the same compromise: generate a short instrumental bed, loop it under the whole cut, duck it under the voiceover, ship. Some of that is the right musical answer. A lot of it is a workaround for tracks that arrived at a length nobody asked for. Once you can ask for the length, the compromise stops being forced and starts being a choice.

    What the release actually contains

    Google's announcement names four improvements (Lyria 3.5 in Flow Music):

    • Improved musicality. Richer melodic structure, less of the wandering that gives a generated track away.
    • Better lyrics. Higher quality and, more usefully, closer adherence to what you asked for and to the structure you specified.
    • More expressive vocals. More emotional range and better pronunciation.
    • Creative control. Direct control over tempo and duration. This is the one that moves work off the timeline and back into the generation.

    Two more capabilities matter for scoring and get less attention than the quality bumps: the model generates tracks up to three minutes, and it takes an image as an input alongside text.

    If you have been evaluating Google's music line generally, the earlier hands-on with Lyria still covers the surface-by-surface picture.

    The loop tax nobody prices in

    Looping a short bed is cheap and it is fine, until you count what it costs you. Four things go wrong, in roughly this order of severity:

    1. The seam. Every loop point is a place where the arrangement resets. A reverb tail gets cut, a cymbal decay disappears, the low end drops for a frame. You can crossfade it away, and most people do, but crossfading a bed four times over ninety seconds slowly smears the whole thing.
    2. No development. A 30-second bed has one idea. A ninety-second video that reuses that idea three times has one idea for ninety seconds. Music that does not go anywhere is the single most reliable tell that a video was assembled rather than made.
    3. Recognition. Viewers detect loops faster than creators expect, because they are not concentrating on the picture the way you are. The second repeat is usually where it lands.
    4. Collision with the voiceover. Loop points ignore your script. When the reset lands mid-sentence, the ear hears the music, not the line. Ducking helps but does not fix a structural event landing in the wrong place. The mechanics of that are in ducking music under a voiceover.

    None of that is an argument to stop using beds. It is an argument to know what you are trading away when you do.

    When a full-length vocal beats a looped bed

    The decision comes down to what the music is doing in the piece, not to how long the piece is.

    Situation Reach for Why
    Video is under 30 seconds, sound-off default Instrumental bed Nothing to develop; the loop never repeats
    Continuous voiceover front to back Instrumental bed A vocal competes with the narration for the same frequency range and the same attention
    Music-led piece with no VO (brand film, montage, launch reel) Full-length vocal The track is carrying the structure, so it needs structure
    The lyrics say something the picture cannot Full-length vocal Only worth it if the words are actually doing work
    Multi-section edit with a build and a payoff Full-length vocal or a long instrumental You need a real arrangement arc, not a repeated bar
    Series content with a fixed sonic identity Bed, plus a signature motif Consistency beats novelty; see sonic branding
    Ad with a hard 15 or 30 second slot Bed cut to length A three-minute arrangement chopped to 15 seconds is worse than a 15-second idea

    The honest summary: a full-length vocal wins when the music is a participant in the piece, and loses when it is furniture. Most brand video wants furniture. Launch films, artist content, and anything you expect people to watch with the sound deliberately on want a participant.

    One caveat that applies to every vocal generation regardless of model: read the words before you spend a render on them. Vocal tracks fail on lyrics far more often than they fail on production, and the fix is a review step, not a better prompt. The workflow for that is in lyrics first.

    Tempo and duration control are the edit-facing features

    Vocals get the headline. Tempo and duration control are the two that change how you actually work, because they move the negotiation from the timeline back to the generation.

    Duration control means you ask for the length of your cut rather than a round number. If the edit locks at 47 seconds, generate 47 seconds. You stop paying the trim tax, where a 60-second track gets cut to 47 and the ending is now a fade instead of an ending.

    Tempo control means you can pick a BPM that agrees with your frame rate. This is worth a minute of arithmetic. At the 25 fps that video defaults to, one beat lasts 1500 / BPM frames. That lands on a whole frame only for tempos that divide 1500 cleanly:

    BPM Frames per beat at 25 fps Cuts land on frame?
    150 10 Yes
    125 12 Yes
    120 12.5 No
    100 15 Yes
    75 20 Yes
    60 25 Yes

    If you are beat-cutting, requesting 100 or 125 instead of 120 means every cut point is an integer frame and your edit stays locked to the music for the whole track rather than drifting by half a frame per beat. Over 180 beats that drift is visible. If you are not beat-cutting, ignore this entirely and pick the tempo that sounds right.

    A prompt shape that uses both controls, rather than describing a mood and hoping:

    Indie-folk, female lead vocal, warm and unhurried.
    100 BPM, 47 seconds, single verse into one chorus, no bridge.
    Acoustic guitar and upright bass, brushed kit entering at the chorus.
    End on a resolved chord, not a fade.
    

    The last line matters more than it looks. "End on a resolved chord" is the difference between a track that finishes and a track your editor has to fade out under the final frame.

    Getting the track into a finished video

    Lyria 3.5 lives on Google's surface, so the practical path is: generate there, bring the audio file in, and do the assembly where the rest of your edit already is.

    Inside Versely, that is a straightforward upload-and-attach. The agent can attach music or a voiceover to a video directly, and the video editor treats the track as one more item on the same timeline as your clips, captions and overlays. Because the editor is EDL-based, swapping the track later re-renders the same timeline rather than forcing a rebuild. Run preview: true for a free 480p pass to check where the music sits against the picture before you commit; it carries a short per-user cooldown, and the final export is charged once no matter how many clips are on the timeline.

    If what you need is the bed rather than the song, Versely's own catalog covers that natively. Suno Sounds V5.5 is listed at 2 credits a generation and ships with looping, tempo and key controls plus lyrics capture, which is the right tool for a short repeating bed where the whole point is that it does not develop. The broader AI music generator is the entry point, and text-to-music covers the terminology if any of this is new.

    FAQ

    Does a three-minute ceiling mean I should stop generating beds?

    No. Beds are the correct answer for most brand and social video, because most brand and social video wants music that stays out of the way. The ceiling matters when the music is meant to be noticed, which is a smaller share of what most teams ship than the excitement around the feature suggests.

    Is Lyria 3.5 available inside Versely?

    Not as a catalog model. The music models listed in the Versely catalog are the Suno Sounds line, and the practical route for a Lyria track is to generate it in Flow Music and bring the audio file into your Versely timeline like any other upload.

    Should I generate one three-minute track or three one-minute tracks for a long edit?

    One track, almost always. Three separate generations will not share a key, a tempo or an arrangement logic, and stitching them produces exactly the seam problem that loops have, only worse because the material either side of the join is different. If you genuinely need sections, generate long and cut, or extend a single track rather than starting fresh, which is the approach in extending music into seamless background tracks.

    Do vocals help or hurt sound-off viewing?

    They hurt, slightly, because a vocal track that nobody hears is a wasted render and the lyric carries information the captions do not repeat. If most of your audience watches muted, spend the effort on captions and on a bed that reads well at low volume rather than on a top line.