The last frame of the source becomes the conditioning state for the new segment, in the same way an input still conditions image-to-video. Motion that was underway carries through, which is why an extension picked at a moment of movement joins more convincingly than one picked on a hard stop.
Drift is the thing to watch. Each extension is conditioned on the previous one's ending, so small deviations in colour, face or set compound. Two extensions usually hold; five in a chain rarely look like the same shot as the original.
It is the honest answer to fixed clip lengths. Most generators cap a single job at a handful of seconds, and extending is how you get past that without stitching two unrelated generations together and hoping the cut hides it.
In practice
- Choose an extension point mid-motion — a static final frame gives the model nothing to continue.
- Extensions inherit the source's resolution and frame shape; you cannot change format mid-chain.
- Colour grade the whole chain at the end, not per segment, or the joins become visible.
Models that extend a clip
Catalog entries that continue an existing video past its last frame. 4 of the 296 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| Flux 3 Extend Video | Black Forest Labs | Video |
| VEO 3.1 Extend Video | Video | |
| Grok Imagine Extend | Grok | Video |
The mistake to avoid
Chaining extensions to reach a target length. Beyond two or three the accumulated drift is usually worse than cutting to a fresh shot.
Go deeper
Runway Video Extend: Stretching Your Best Takes
A working guide to Runway Video Extend: when stretching a take beats regenerating, continuation prompts, drift limits, and the cost math for editors.
Where you will run into it
- Extend a Video's Length — Your clip ends too soon. Keep it going.
Related terms
First-last frame
First-last frame generation takes two stills — where the clip starts and where it ends — and generates the motion that gets from one to the other.
Temporal consistency
Temporal consistency is how well a generated clip keeps things the same from one frame to the next — a shirt that stays the same colour, a background that stays put.
Duration
Duration is how long a generated clip runs, chosen before generation from whatever lengths the model supports rather than trimmed afterwards.
Image-to-video
Image-to-video animates a still you supply: the picture becomes the opening frame, and the prompt describes only what happens next.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.