Industry

    The 30-Second Single Take: AI Video's Duration Race

    Seedance 2.5 and Wan 3.0 chase 30-second single takes, Flux 3 ships 20s with audio. Why duration is hard to ship, and how to hit 30s on Versely today.

    Versely Team7 min read

    Twelve months ago, five to eight seconds was what most AI video models could hold in one generation. That ceiling has become the single most contested spec in the category this year — more contested than resolution, more than prompt adherence — because duration is the one capability that can't be faked with a bigger GPU budget alone. A model has to hold an entire sequence coherent at once, and that requirement compounds with every additional second rather than scaling with it.

    The confirmed 30-second claim, and the one that isn't yet

    Alibaba's Wan 3.0 opened public beta earlier this month, and the headline spec is a genuine break from the pack: a continuous clip up to 30 seconds in a single pass, with no stitching. That's live in beta today, and it's the clearest evidence yet that a 30-second ceiling is reachable with current architectures rather than being a marketing round number.

    Seedance 2.5 gets cited in the same breath, and it's worth being precise about why that citation is softer. As of this writing, ByteDance hasn't published Seedance 2.5 specifications — the model isn't released. Tracking pages built on industry leaks point to a longer ceiling than Seedance 2.0's current 15 seconds, and native audio is expected to carry over, but the specific number circulating as "30 seconds" isn't in ByteDance's own materials yet. Worth knowing before you plan a shoot around it: "claimed" and "shipped" are different categories in this race, and mixing them up is how a production calendar ends up built on a leak instead of a spec sheet.

    The model that actually delivers 30 seconds today — shipped, billable, no beta waitlist — is Grok Imagine Video on Versely. It runs the full range from 6 up to 30 seconds in a single generation, at 7 credits per second for HD output (5 for SD) — a full 30-second HD take runs 210 credits, one generation, no stitching. If the 30-second single take is the thing you need this week rather than the thing you're watching for, this is where it already exists.

    The 20-second and 15-second tier

    Below the 30-second claims sits a more crowded, more confirmed tier. Black Forest Labs' FLUX 3 Video generates clips up to 20 seconds with native audio — dialogue, sound effects, and ambient sound generated alongside the picture, not layered on after. MiniMax H3, released July 31, tops out at 15 seconds at 2K with native stereo audio, and Kling 3.0 Turbo caps at the same 15 seconds — a number Versely's own Kling 3 Turbo listing matches exactly, running 3 to 15 seconds per generation. Twelve months ago, 15 seconds would have been the headline spec of the year. In this field it's the middle of the pack.

    Why duration is the hardest spec to ship

    Duration is a menu, not a slider, for a structural reason: the model has to hold the whole sequence in mind at once to keep it coherent, and that cost grows with length rather than scaling linearly against it. Temporal consistency — the thing that keeps a shirt the same color and a background from crawling frame to frame — is already the hardest unsolved problem in video generation at eight seconds. Stretching the same architecture to thirty asks it to hold that stability nearly four times as long, which is why the models claiming it are the newest, largest releases in the category rather than incremental updates to existing ones.

    The credit math on Versely's own catalog shows the cost curve directly, because several models bill duration and resolution as separate multipliers on the same per-second rate. LTX 2.3's fast tier runs 4 credits per second at SD/HD and 16 credits per second at 4K — the same clip, the same length, four times the price for resolution alone, before duration even enters the equation. Multiply that resolution premium by a duration that's also climbing and the two variables compound: a 20-second LTX 2.3 clip at 4K runs 320 credits against 80 for the same 20 seconds at HD. Every model chasing a longer ceiling is quietly making the same bet — that buyers will pay a steeper curve for a single continuous take than they would for the equivalent length assembled from cheaper, shorter pieces.

    What a longer single take replaces

    The honest previous answer to "I need more than my model's ceiling" was one of two moves: extend from the last frame, or generate several shots and cut them together. Assembly usually wins on quality, because a cut hides the seam an extension has to survive — but it means directing multiple generations instead of one. A native 30-second take collapses that decision for the specific case of a single continuous shot: there's no seam to hide, because there's no join.

    Extension isn't going away, though, and the cost math explains why. VEO 3.1 Extend pushes a clip forward in fixed 7-second hops — 140 credits per hop without audio, 280 with — so reaching 30 seconds from an 8-second base takes three or four extend calls stacked on top of the original generation. That's a legitimate route on a model that doesn't natively reach 30 seconds, but it's a different cost shape than one native generation: you're paying per hop, and each hop only sees the last few seconds of the previous one as its reference, so drift is cumulative in a way a true single take doesn't have to worry about.

    Three ways to a 30-second spot on Versely, with the credit math

    Native, cheapest path — Grok Imagine Video. One generation, 6–30 second range, 210 credits for the full 30 seconds at HD, 150 at SD. No extends, no stitching, one prompt covering the whole arc.

    Native with audio, then extended — FLUX 3. FLUX 3 Text to Video generates with native dialogue and sound design built in, at 9 credits per second — 43 to 170 credits across its 5–20 second range. To go past 20 seconds, FLUX 3 Extend continues an existing clip beyond its final frame at 21 credits per second, but the source clip has to be under 15 seconds and under 50MB. The working sequence: generate a base clip comfortably under that 15-second cap, then extend it by up to 20 more seconds — 103 to 410 credits — for a continuous FLUX take past 30 seconds total, native audio carried through both generations.

    Extend chain — anything else. For a model whose native ceiling is well under 30 seconds and has no matching Extend variant of its own, plan the shot as extend hops from the start. Know your per-hop cost and count before generating the base clip, because that number multiplies fast — on VEO 3.1, reaching 30 seconds from an 8-second base is three or four extend calls, not one all-at-once generation.

    A practical note for the Grok Imagine route: write the prompt as one continuous arc — establishing beat, the turn or reveal, the closing beat — rather than three separate ideas and hoping the model paces them evenly across the full 30 seconds. Single continuous takes read the whole prompt as one shot; front-load the structure the way you would for a single long take on a real set, because that's functionally what it is.

    What still degrades past twenty seconds

    None of this makes thirty seconds free of the problem it was always going to have. The longer a single take runs, the more a busy background, fine repeating detail, or a lot of independent motion has to stay locked down the whole way through — the failure mode is the same one that shows up at eight seconds, just with more runway for it to compound. A clean, simple frame holds up at thirty seconds far more reliably than a crowded one does, which is the same advice that's always applied to duration, just mattering roughly four times as much now that the ceiling has moved.