Guides

    24fps clips on a 25fps timeline

    Video models render at 24fps and Versely timelines default to 25. The three ways to conform, the arithmetic behind each, and the one that wrecks lip sync.

    Versely Team9 min read

    A one-frame-per-second difference sounds like the kind of thing you can ignore. It is, right up until someone speaks on camera.

    Here is the mismatch, stated as plainly as it deserves: every video model in the catalog that exposes a frame rate at all offers 24fps, most add 30 and some add 60 — and not one of them offers 25. Versely's default video frame rate is 25. So the moment a generated clip lands on a sequence, something has to give, and the something depends on a choice you make — or, if you don't make it, a choice made for you. There are exactly three ways out of this and one of them quietly ruins any clip with dialogue in it.

    A large monitor showing video editing software with multiple timeline tracks

    Where the mismatch comes from

    There is no universal frame rate. 24 is the cinema convention, 25 is the PAL broadcast convention, and both are entirely reasonable defaults that happen to disagree by about four percent. Generative video models inherit a rate from whatever their training and inference pipeline settled on, which is why the models in any large catalog do not all agree with each other.

    The practical consequence is that frame rate is a per-model property you have to know, in the same way you already know a model's duration ceiling and aspect ratios. Veo 3.1 exposes 24, 30 and 60, Kling 2.5 Turbo exposes the same three, and plenty of models expose no choice at all. What none of them expose is 25, so a 24fps render against a 25fps sequence is a four-percent problem you will meet repeatedly — and four percent is small enough to slip past a casual review and large enough to be obvious on a face.

    If your clips disagree with each other rather than with the timeline, that is a different article. This one is about the conform you still owe once the whole set agrees.

    The arithmetic is worth having in front of you, because every option below is a different answer to the same equation:

    • 25 ÷ 24 = 1.041666…, so a 24fps clip played at 25fps runs 4.17% fast.
    • A 10-second clip becomes 9.6 seconds.
    • Picture and audio, if they separate, drift by 40 milliseconds per second — 400ms over ten seconds.

    Forty milliseconds a second is the number to remember. It's below the threshold where a single frame looks wrong and well above the threshold where accumulated drift looks wrong.

    Option 1: retime the clip

    Retiming means letting the 24 frames play out over 25 frame slots. You speed the clip up by 4.17% and the frame counts reconcile themselves.

    This is the option that costs you lip sync, and it does so in one of three ways depending on how the audio is handled:

    Picture and audio retimed together, pitch following. Sync survives. Pitch rises by about three quarters of a semitone (12 × log₂(25/24) ≈ 0.71). On ambience and most music beds nobody notices. On a voice the audience knows, or a voice you've matched to a brand, it reads as subtly wrong in a way people describe as "off" without being able to name it.

    Picture and audio retimed together, pitch corrected. Sync survives, pitch survives, and you pay in time-stretch artefacts. Sibilants and transients smear. Four percent is a small stretch, so this is often acceptable — but it is a processing pass on your audio that you did not want.

    Picture retimed, audio left alone. Sync is destroyed. This is the default outcome when a conform happens implicitly rather than deliberately, and it's the worst of the three. The drift is 40ms per second, so lips and words are visibly apart by the second sentence and comically apart by the end of a fifteen-second clip.

    Retiming is a legitimate choice when the clip has no speech and no music that needs to land on a beat. For anything with a mouth moving in it, pick something else. The editor handles playback speed changes as a described instruction rather than a numeric field, so if you do go this route, state explicitly whether pitch should follow or be held.

    Option 2: duplicate one frame a second

    24 frames into 25 slots needs one extra frame per second. Duplicate one and you're done.

    What this buys you is the thing retiming gives away: audio is untouched, duration is untouched, sync is untouched. A 10-second clip stays 10 seconds. A voiceover stays aligned. Captions built against the original timing stay valid, which matters more than it sounds if you've already run forced alignment to time them.

    What it costs is a once-per-second hitch, because one frame holds for two slots instead of one. Whether that reads depends entirely on what's moving:

    • Locked-off shots, talking heads, slow pushes: invisible. Genuinely invisible.
    • Moderate camera moves: detectable if you're looking for it, not otherwise.
    • Fast horizontal pans, whip moves, anything crossing frame quickly: visible as a stutter, and once you've seen it you can't unsee it.

    The refinement for that last row is frame interpolation instead of duplication: synthesise the missing frame from its neighbours rather than repeating one. Motion stays smooth, at the cost of an extra processing pass and the risk of interpolation artefacts on fast or complex motion. For most sequences duplication is the right call and interpolation is over-engineering.

    Option 3: standardise the render rate

    The third option is to stop generating the problem. Instead of conforming at the timeline, decide the rate at the source and stick to it for the whole sequence.

    In practice this means routing every shot in a sequence through models that share a render rate, rather than picking per-shot on quality and discovering the mismatch in the edit. It is the only option with zero artefacts, because there is nothing to conform.

    The reason to take it seriously is not really the artefacts. It's consistency. A sequence where shot one was retimed, shot two was frame-duplicated and shot three needed nothing has three different motion cadences in it, and that inconsistency is more noticeable than any single artefact. Mixed-rate sequences are the actual failure mode. A uniform 24fps sequence conformed once, the same way, throughout, looks like a deliberate choice. A sequence conformed shot by shot looks like a mistake, because it is one.

    The cost is real: you're constraining model choice on a technical property rather than picking the best model for each shot. For a one-off clip that's a bad trade. For a six-shot series that has to feel like one piece, it's usually the right one.

    Picking one, and checking it before you pay

    The decision collapses to what's on the audio track.

    What's in the clip Conform
    Dialogue, voiceover, anything lip-synced Duplicate a frame
    Fast pans or whip motion, with dialogue Interpolate
    Ambience or music only, no speech Retime — it's the cleanest picture
    Music that has to land on a beat Duplicate a frame
    A multi-shot series that must match Standardise at the source

    Then check it before you spend on the finished render. The editor is EDL-based, so the timeline is a re-renderable object rather than a one-shot export, and preview: true gives you a 480p pass at no credit cost with a short per-user cooldown between them. The final export is charged once no matter how many clips are on the timeline, so the preview genuinely is the cheap check rather than a token gesture.

    Both failure modes in this article survive the downscale. A sync break is an audio-versus-picture problem, so 480p shows it perfectly — you don't need resolution to see lips disagreeing with words. Frame-duplication judder is a temporal artefact, so it reads intact too. That is not true of every post decision, which makes this one of the cases where the 480p pass, cooldown and all, genuinely settles the question. Watch the whole clip rather than scrubbing, and watch it with sound on.

    The cost side is laid out in the previews and final export breakdown if you want the exact shape of what gets charged when.

    FAQ

    Why do video models render at 24fps and not 25?

    24fps is the cinema convention, and models aiming at a filmic look inherit it. 25 is the PAL broadcast convention. Neither is a defect — they are two reasonable traditions that disagree, and the models in the catalog landed on the cinema side of it. That is the entire reason a conform step exists. Knowing a model's render rate up front belongs alongside knowing its duration ceiling and aspect ratios.

    Does a 4% speed change really break lip sync?

    If picture and audio are retimed together, no — they stay aligned, and you deal with a small pitch shift or a small stretch artefact instead. If only the picture is retimed, yes, badly: the drift accumulates at 40 milliseconds per second, so it's clearly wrong within a couple of seconds and unusable across a full clip. The danger is that the picture-only version is what happens when nobody makes the choice deliberately.

    Is frame duplication or interpolation better?

    Duplication for most content, because it's simple, predictable and artefact-free apart from a once-a-second hold that locked-off and slow-moving shots hide completely. Interpolation for fast horizontal motion, where that hold becomes a visible stutter. Interpolation costs an extra pass and can smear complex motion, so it's the exception rather than the default.

    Can I mix 24fps and 25fps clips in one sequence?

    You can, and it works, but conform every clip the same way rather than shot by shot. Mixed conform methods within one sequence produce different motion cadences between shots, and that inconsistency is more noticeable to a viewer than any single conform artefact. If the series has to feel like one piece, standardise the render rate at the source instead.

    Assemble and check the whole thing in the AI video editor, which is also where the merge step lives if all you need is several generations stitched into one cut.