Guides

    Continuation Prompting: Feeding a Model Its Own Last Four Seconds

    Continuation prompting explained: hand a model the end of an existing clip plus an instruction for what happens next, and extension becomes direction.

    Versely Team7 min read

    Worth separating from the start, because the two get confused constantly: first-last-frame generation interpolates between two static images you already have, filling in a transition the model has to invent from scratch on both ends. Continuation prompting is a different job entirely — it hands the model a slice of video that already happened, momentum and all, and asks it to keep going. One invents a bridge between two fixed points. The other extends a path that's already in motion.

    Extension is not just "make it longer"

    The lazy framing of "extend this video" treats it as a duration problem — the clip is eight seconds and you need twelve, so bolt four more seconds on the end. That framing misses what actually makes continuation prompting useful: because the model is conditioned on real footage right up to the cut point, the instruction you give it isn't "generate more video," it's "here's where we are — now direct what happens next." A generic text-to-video prompt has to establish a scene from nothing. A continuation prompt inherits an established scene and only has to answer one question: what happens now. That's a meaningfully easier, more precise creative instrument than it looks like from the outside, and it's why the instruction that goes with it reads more like a director's line than a generation prompt.

    What actually crosses the seam

    The detail that makes this work is what the model is allowed to carry across the cut. Black Forest Labs' own description of FLUX.3's video continuation feature is specific about this: the model takes up to four seconds of existing video and audio and is told what should happen next, and it accounts for both components to continue "movement, camera behavior, dialogue, and audio across the video seam." That's four distinct threads staying alive through the cut — not just visual pixels extending forward, but the camera's actual motion, whatever a character was saying, and the audio bed underneath all of it. A continuation that only matched pixels and dropped the audio, or kept the audio but reset the camera to static, would announce itself as a splice immediately. Carrying all four through the seam is what makes it read as one continuous shot instead of two clips stitched together.

    The catalog options, and what each is built for

    Versely's extend-video catalog spans a few different engines, and they're not interchangeable on capability:

    • Flux 3 Extend Video (21 credits) continues an existing clip beyond its final frame, matching the source's motion and scene — the source clip needs to be an MP4 under 50MB and under 15 seconds, which makes this the tool for extending a recent generation rather than an old archival file.
    • LTX 2.3 Extend Video (10 credits) continues motion or action past the current end — or before the current start, which is worth knowing since most continuation thinking only goes forward. It supports extensions up to a generous 20 seconds, the longest ceiling in the catalog for this job, useful when a single continuation needs to cover a lot of narrative ground rather than a quick few-second beat.
    • VEO 3.1 Extend Video (40 credits) and VEO 3.1 Fast Extend Video (15 credits) extend duration on Google's model, with the Fast variant trading some of the flagship's ceiling for a cheaper, quicker pass — a reasonable split for exploring a continuation direction cheaply before committing the full-price version to the shot that's actually keeping.

    None of these are strictly "better" in the abstract — the right pick depends on how long a continuation the shot needs, whether the source clip is fresh enough to fit a size/duration ceiling, and how much the specific scene leans on dialogue and camera work surviving the cut versus a simpler continuation of ambient motion.

    Directing the continuation

    Because the model already has the scene, the instruction that gets the best result reads less like a full scene description and more like a director calling out what happens next mid-take. A few patterns that work better than a generic "continue the video":

    • Name the camera behavior explicitly, even if it's "hold." A continuation prompt that says nothing about the camera risks the model reinterpreting the existing motion rather than extending it — "keep the slow push-in going" carries the established move forward on purpose.
    • Give the next beat, not a new scene. "The character turns and walks toward the door" continues what's already happening; "cut to a wide shot of the street outside" is asking for a scene the model has no established footage for, which defeats the point of feeding it a continuation in the first place.
    • Keep dialogue continuations short and specific. If a character was mid-sentence, give the next line directly rather than a general topic — the model is trying to carry vocal performance and audio timing across the seam, and a vague instruction gives it more room to drift from the established voice and pacing.
    • State what should change, since everything else is assumed to hold. The model's default is continuity; the instruction's job is to specify the one thing that's different now, not to re-describe everything that's staying the same.

    A Versely walkthrough

    A continuation request in chat, grounded in a real generated clip:

    "Extend this video by another 6 seconds — keep the same slow camera push-in, and have the character glance toward the door as if they just heard something."

    That routes to extend a video, which takes the existing clip, generates additional footage continuing from where it ends, and returns a longer version with the new footage picking up the established motion rather than resetting it — the mechanics of the underlying tool, including which models are currently live, are laid out on extend a video's length if the chat shorthand above needs unpacking first. Model choice matters here more than in a fresh generation, since the source clip's length and file size can rule options in or out before creative fit even enters the decision — Versely's best video extender models ranking is the place to check which engine is currently ahead before defaulting to whichever one was used last.

    One honest limit worth planning around: a continuation is new generated footage, not a slow-motion stretch of the original frames, and the join isn't guaranteed to be invisible on every attempt. Review the seam itself — the exact frame where original footage hands off to generated continuation — before treating an extended clip as finished, the same way any video extend job deserves a specific check rather than a glance at the finished cut.

    FAQ

    How is continuation prompting different from first-last-frame generation?

    First-last-frame interpolates a transition between two static images you supply, inventing the motion between them from scratch. Continuation prompting extends from real, already-generated footage, carrying its established motion, camera behavior and audio forward rather than bridging between two fixed endpoints.

    Does the extended footage keep the original audio going too?

    That's the point of feeding the model both video and audio as source material — a continuation that only matches pixels and drops the audio, or resets the camera to static, reads as an obvious splice. The stronger continuation models are built to carry movement, camera behavior, dialogue and audio together across the cut.

    How long a clip can I extend in one pass?

    It varies by model — LTX 2.3 Extend Video supports extensions up to 20 seconds, the longest ceiling currently in the catalog, while other engines are built for shorter, more targeted continuations. Check the specific model's limits against how much narrative ground the continuation actually needs to cover.

    What should I actually write in a continuation prompt?

    Treat it like a director's line mid-take, not a fresh scene description: name what the camera should keep doing, give the next narrative beat rather than a new scene, and state only what's changing — the model already has everything else from the source clip.

    Pick a recent clip that ended too soon, and ask Versely to extend a video with a specific next beat instead of a generic "continue" — direction is the whole trick.