Guides

    Flux 3 Extend: Continuing Shots Without Re-Rolling

    How Flux 3 Extend Video continues an existing clip's motion and audio past its final frame, and when it beats first/last-frame stitching for longer takes.

    Versely Team9 min read

    Every text-to-video model has the same tell once you push it past its comfort zone. Ask for a beat that happens at second fourteen of an eight-second generation and most models don't extend the shot — they re-imagine the whole thing. Wardrobe drifts, framing resets, the take you actually liked is gone, replaced by a cousin of it. Flux 3 Extend Video exists specifically to route around that failure mode. It doesn't regenerate your clip with a longer duration slider; it picks up the real clip — motion, room tone and all — at its last frame and keeps going.

    Of the four modes in Black Forest Labs' Flux 3 family, Extend is the one nobody writes about. Text-to-video gets the reviews, first/last frame gets the transitions-and-loops posts, image-to-video gets the "animate this photo" tutorials. Extend gets skipped because you only reach for it after you already have a shot worth keeping — which makes it the least demoed and most useful mode once you actually need it.

    A glowing computer monitor in a dim room showing video editing software with timeline markers

    What Extend actually does

    Black Forest Labs unveiled Flux 3 on July 23, 2026 as a single multimodal model spanning image, video and audio rather than three separate products bolted together. On the video side specifically, BFL's own launch post lists four modes: text-to-video, image-to-video, keyframe (first/last frame) generation, and generative continuation from an input video and its audio. That last one is Extend, and the phrase "from an input video and its audio" is the whole point — it's not a text-to-video call with your old clip as a mood board, it's a continuation of a specific audiovisual state that already exists.

    You can see the difference in the input shape alone. Text-to-video, image-to-video and first/last frame all start from a prompt, a still, or a pair of stills — something you write or compose. Extend is the only mode in the family that requires an actual video upload: one MP4 source clip, and nothing else gets it started. On Versely, that source clip has real limits worth knowing before you reach for it — under 50MB, under 15 seconds, MP4 only — because Extend is built to continue a hero shot, not to process an arbitrary upload.

    It's also priced like a distinct move rather than a default one. Flux 3 Extend Video runs 21 credits per generation against 9 credits for Flux 3 Text to Video, Flux 3 Image to Video or Flux 3 First Last Frame to Video. That gap is a useful signal: Extend isn't the cheap way to pad a clip to a longer duration, it's the tool for the specific case where re-rolling would throw away a take that's already working.

    Extend vs. first/last-frame stitching

    These two modes get confused because both produce "more video from what you already have," but they solve opposite problems.

    First/last frame takes two stills — where the clip starts and where it ends — and generates the motion connecting them. That's a bounded problem: you're telling the model exactly where the shot has to land, which is what makes it the right tool for loops (same still in both slots) and transitions (the last frame of shot one becomes the first frame of shot two). The glossary entry on first-last frame covers the mechanics in more depth — the short version is that you own the destination, and the model finds a route there.

    Extend has no destination. It has a source clip and a prompt describing what happens next, and it reads the clip's own final frame as the jumping-off point rather than a fixed endpoint you supplied. That's the right tool when you don't know — or don't want to pre-decide — where the shot ends, you just want more of the motion that's already underway. Trying to force that job through first/last frame means manufacturing an end frame you don't actually have yet, which means deciding the ending before you've generated it. Extend skips that step entirely.

    The rule that actually holds up in practice: if you know exactly where a shot needs to land — a seamless loop, a match-cut into a second scene — use first/last frame. If you just have a good take that stopped too soon, use Extend. The Extend video glossary entry frames this the same way: video extend is "the honest answer to fixed clip lengths," for when stitching two unrelated generations together and hoping the cut hides it isn't good enough.

    The seam that actually matters is audio, not video

    Flux 3 generates video with native audio up to 20 seconds long — a first for Black Forest Labs, and that native-audio detail is what makes Extend meaningfully different from just cutting between two separately generated clips in an editor.

    Picture the alternative: you generate an 8-second clip, like the result, then generate a second 12-second clip meant to follow it, and cut them together on a timeline. Both clips have native audio — but each one's audio was generated as its own independent event. Room tone, ambient noise, even the character of silence between sounds, all sampled fresh, twice. The picture cut can be invisible if the framing matches. The audio cut almost never is, because two independently generated ambiences don't share a single consistent room even when they're describing the same room.

    Extend doesn't have that problem because there's only one audio-generation event that keeps running. Since the continuation mode works from the input video's audio directly, not a fresh audio generation prompted by a text description of the scene, the room tone that was present in your source clip is the room tone the extension continues — no second sampling, no seam to hide. For anything where continuous ambience actually matters — a kitchen scene, a street exterior, dialogue in a specific space — that's the real argument for reaching for Extend over generating two clips and hoping the cut disguises the difference.

    Walkthrough: an 8-second hero shot into a continuous ~20-second take

    This is the pattern for turning one good short clip into a longer, seamless one without re-rolling the whole thing.

    1. Generate your hero shot first, with Flux 3 Text to Video or Flux 3 Image to Video, and land the last second or two of it mid-motion rather than on a hard stop — someone still walking, water still pouring, a hand still reaching. A generation that settles into stillness gives Extend nothing to carry forward; the model needs momentum to continue, not a resting pose.
    2. Export that clip. An 8-second hero shot clears Extend's source-clip rules with room to spare — MP4, well under the 50MB ceiling, well under the 15-second cap.
    3. Open Flux 3 Extend Video and upload the source. Prompt only the new beat — what happens after the frame you're extending from — rather than re-describing what's already in the shot. If there's dialogue in the source, write the next line in the same voice, not a fresh take on the first one.
    4. Set the duration. The selector runs 5 to 20 seconds; request 12 seconds and, placed after your 8-second source, that's a continuous ~20-second take. Match resolution and aspect ratio to your source clip so the join doesn't visibly change lens partway through.
    5. Render and check the handoff frame specifically — both the motion continuity and the audio continuity — then place the two clips back-to-back on the timeline.
    6. Don't extend the extension. Treat one Extend pass as the move, not the first link in a chain. Each additional pass is conditioned on the previous pass's own ending, not your original footage, so drift compounds fastest exactly where you're least likely to be watching for it — the audio bed. If a sequence needs more than about 20 seconds on one beat, that's usually a sign it wants a second shot, not a longer first one.

    What the source-clip limits are actually telling you

    The 15-second, 50MB, MP4-only ceiling on the source clip isn't an arbitrary technical constraint — it's a statement about what Extend is for. It can't take a full minute of raw footage and continue it indefinitely; it's built around a short, deliberately chosen hero shot, the same way a director picks one take to print rather than feeding the whole reel forward. The single-video-input limit reinforces this: Extend continues one clip, it doesn't composite several sources into a new one. If you're reaching for it, you should already have the shot you're happy with — Extend's job is to make it longer, not to find it for you. The extend video length guide covers the broader editing pattern if you're stitching multiple extended segments into a finished timeline.

    FAQ

    Does Extend keep the source clip's actual audio going, or generate new ambience that just sounds similar?

    It continues from the input clip's audio directly rather than generating a fresh soundscape from a text description of the scene. That's the mechanism that avoids an audible seam — there's one continuous audio state, not two independently sampled ones meeting at a cut.

    Does the source clip have to come from Flux 3 itself?

    No — the requirement is a file spec, not a same-model rule. Any MP4 under 50MB and under 15 seconds can go into the source slot, whether it came from Flux 3, another model, or a real camera.

    Is Extend the cheap way to get a longer clip?

    No, and that's by design. At 21 credits against 9 for a fresh Flux 3 Text to Video generation, Extend is priced as the deliberate move for a take worth continuing — not the default way to pad a clip to a longer duration.

    When should I use first/last frame instead of Extend?

    When you already know where the shot needs to end — a seamless loop, or a transition into a second scene you've already got a frame for. Extend has no end-frame input; it only takes a source clip and a prompt for what happens next.

    Flux 3 Extend Video is live now inside Flux 3 Text to Video and its sibling modes in Versely — generate the hero shot, keep the take you like, and extend it instead of rolling the dice again.