Workflows

    Twenty-Second Takes Change the Edit, Not Just the Render

    AI video durations are pushing past 20 seconds. That changes where editorial decisions happen, and quadruples the cost of a flaw buried mid-take.

    Versely Team8 min read

    For most of AI video's short history, duration has been a spec you compare on a pricing page — 5 seconds versus 8, versus 10. Once single generations reliably stretch past 20 seconds, duration stops being a spec and becomes an editorial decision, because a 20-second take isn't a longer version of a 5-second take. It's a different kind of shot, made for a different kind of editing, with a different failure economics attached to it.

    Monitor showing a video editing timeline

    What actually shipped

    Black Forest Labs' FLUX 3 Video generates clips up to 20 seconds in HD, with Full HD available through upscaling — shipped via the BFL API and select partners. It's a genuinely new release outside what Versely currently carries, worth naming plainly rather than implying otherwise.

    The good news for anyone who wants to test what this post is actually about today: Versely's own catalog already covers this ground. LTX 2.3 Text to Video Fast supports durations up to 20 seconds at resolutions as high as 4K, and Grok Imagine's image-to-video model lists configurable duration up to 30 seconds. This isn't one lab's flex anymore — it's a trend line across the field. Five to eight seconds was the assumed ceiling for a long stretch of this industry's short history; 20-to-30-second single takes are now something you can point a live, carried model at, not a demo reel promise.

    Why longer changes the edit, not just the render

    At 5–8 seconds, a generation is basically a shot. You cut to it and cut away from it, and the generation itself defines the shot boundary — your editorial judgment lives in the gaps between generations, in sequencing and pacing and the rhythm of the cuts.

    At 20-plus seconds, a single generation can hold multiple beats inside one continuous take: an establishing moment, a turn, a reaction, a change in energy. That relocates a chunk of editorial decision-making from the cut into the prompt itself. You're no longer just sequencing discrete shots for that stretch of the video — you're directing a single unbroken performance and hoping the model sustains it, which is a genuinely different skill than picking good cut points between short generations.

    It also changes which editing tools sit at the center of your workflow. A long single take invites you to cut into it after the fact — fixing or reworking a window inside an existing generation — rather than only appending fresh generations after it. Tools built around extending a video's length and repairing a segment in place become more central than tools built purely around sequencing separate clips.

    The quadrupled cost of a bad take

    Here's the part that doesn't show up on a duration spec sheet. If a 5-second generation has a blown detail — a warped hand at second 3, a continuity break — fixing it costs one regeneration of that whole clip. A 20-second generation with the same category of flaw at, say, second 14 costs a full regeneration of all 20 seconds to fix a single broken window, because you can't cheaply re-roll just the damaged three seconds inside one continuous generation the way you can re-roll an entire short clip.

    Roughly a fourfold increase in length — five seconds to twenty — becomes a roughly fourfold increase in cost for the exact same category of failure, because a flaw anywhere in the take invalidates the whole take, not just the segment where it happened. That reframes risk in a way that matters for planning a shoot list: longer single takes raise the stakes of everything that can go wrong mid-generation — drift, identity loss, an object glitching for a frame or two — simply because there's more running time for something to break, and no way to salvage the good seconds surrounding a bad one within a single generation call.

    The direct mitigation is segment-level repair rather than full regeneration. A retake-style tool that takes a source video plus a start time and duration and regenerates only that window — leaving the rest of the take untouched — is exactly the tool built for this specific economics problem. Reach for it the moment a flaw shows up inside an otherwise-good long take, instead of defaulting to a full re-roll.

    When 20 seconds is the wrong choice, on purpose

    None of this is an argument that longer is always better — it's an argument that duration should follow from what job the shot is doing. Hooks, ads, and fast-paced short-form cuts generally want many short takes stitched with deliberate cut rhythm; a single unbroken 20-second take removes the cut points that make that pacing possible in the first place. Longer single takes earn their place for an uninterrupted product reveal, an establishing or atmosphere shot that benefits from not being cut, a performance beat that loses something if it's chopped up, or b-roll you intend to trim down selectively in post rather than use whole.

    The practical rule of thumb: ask whether the shot wants to breathe uncut, or whether the edit wants material to cut between. Those are different jobs, and the duration you generate at should follow from which one you're actually doing — not from the assumption that a longer single take is automatically the more impressive result. Our guide on choosing a model for long-form output and the duration glossary entry both cover the trade-offs worth checking before you commit a full render to either approach.

    A Versely walkthrough: testing a long take without eating the full cost of a miss

    The workflow that respects the cost math above is straightforward:

    1. Lock the prompt short first. Generate at 6–8 seconds with LTX 2.3 Text to Video Fast to confirm the subject, style, and motion are holding up before committing to a full-length take. This is the cheap iteration pass on the part most likely to need adjustment.
    2. Commit to the full duration once the short version holds. Regenerate the same prompt at up to 20 seconds — you're now spending the longer render on a prompt you've already validated, not one you're still debugging.
    3. Patch, don't regenerate, if something breaks mid-take. If a specific window glitches — a hand at second 14, a background object that flickers — ask the agent directly: "Retake this video from second 12 to second 16, fixing [the issue], and keep the rest of the clip unchanged." That targets just the damaged segment instead of paying for a full 20-second regeneration to fix three broken seconds, which is the exact mitigation for the cost problem this post is describing.

    Browse the full models catalog to compare which of the current long-duration options fits your shot before you commit to a full-length render.

    FAQ

    Why does a longer AI video generation cost more to fix when something goes wrong?

    Because most generation and retake tooling works on the take as a whole. A flaw anywhere inside a single continuous generation typically requires regenerating the entire take to fix it, so a 20-second generation with one bad window costs roughly four times what a 5-second generation with the same flaw would cost to fix — unless you use segment-level repair instead of a full regeneration.

    Does Versely support 20-second video generations?

    Yes. LTX 2.3 Text to Video Fast supports durations up to 20 seconds at resolutions up to 4K, and Grok Imagine's image-to-video model supports configurable duration up to 30 seconds — both live in Versely's catalog today.

    Is FLUX 3 Video available on Versely?

    Not currently. FLUX 3 Video is Black Forest Labs' newest release, shipped via the BFL API and select partners, and sits outside what Versely's catalog carries right now.

    Should I always generate the longest duration a model supports?

    No. Longer single takes suit shots meant to hold uncut — reveals, establishing shots, unbroken performance beats. Fast-paced editing that depends on frequent cuts is usually better served by several shorter generations stitched together, since a single long take removes the cut points that rhythm depends on.

    How do I fix a flaw in the middle of a long AI video take without regenerating the whole clip?

    Use a segment-level retake tool that takes a start time and duration and regenerates only that window of an existing video. Asking for a targeted fix on just the broken seconds avoids paying for a full regeneration of an otherwise-good long take.

    Match the duration to the job the shot is doing, then protect the render with segment-level fixes instead of a full re-roll. Start testing long takes in the models catalog.