Guides

    extract_video_frames pulls stills; it does not render new ones

    Frame extraction is one named processing job. Using it as analysis, upscale, or a fresh generate spends a tool on the wrong output.

    Versely Team4 min read

    extract_video_frames pulls stills from a clip you already have. It does not upscale them. It does not restyle them. It does not describe the video beat by beat. It does not invent a better thumbnail than the frame that exists. Give it a prompt instead of a video_url and there is nothing to grab.

    Extract frames from a video is the named agent job. Grab exact stills from any point in a clip. Required: video_url. Optional: start time, end time, frame count, fps, output format. Output is image URLs. Cost is a lightweight processing step, not a generative model — check your plan. That cheapness is why people run it when they meant a generate, and why people skip it when they meant a salvage.

    A frame from a failed take is often worth more than a reroll. Extraction exists so you can keep the one second that worked without buying the other seven again.

    What the job is

    Thumbnails. Reference stills for a later edit or image-to-video. A mid-clip hold you want as a poster. You may chain a grabbed frame into edit a photo or turn a photo into a video. Those chain steps are new jobs with new confirms. Extraction only mints the stills.

    If you cannot tell which frame to grab, you may want a structured read of the clip first: get a detailed breakdown of a video (analyze_video) extracts sample frames and describes beats, on-screen text, optional transcript. That is analysis, priced as analysis. Do not call it when you already know "the 5-second mark."

    What it is not

    • A new image generate. generate_images invents a thumbnail. Extraction copies a moment that already survived motion. If the take is close, grab. If the take never had the composition, generate — do not sample twelve junk frames hoping one is a poster.
    • Upscale. Upscale an image is a different model. A 720p grab stays 720p until you pay that row.
    • Prompt-match QC. Get AI feedback on a generation is review_generation. It tells you whether the clip matched the brief. It does not return a still.
    • Virality or trend analysis. Scoring a public post is analyze_trend. Scoring your unpublished file against a prompt is review. Pulling JPEG URLs is this.

    Rerolling a whole video because you needed a thumbnail is the expensive substitute. Extract first. Decide whether the still is enough.

    The test

    Do you already know the clip, and do you need one or more frames from it as images?

    Yes: this job. Name the time if you have it. Pull the URLs.

    If you need to understand the clip, analyze. If you need a different picture, generate or edit. If you need the picture bigger, upscale the grab. Do not ask extraction to become those tools.

    FAQ

    How is this different from the frames analyze_video returns?

    analyze_video is a breakdown: style, beats, text, optional speech, plus sample frames. extract_video_frames is only the grab, with time and count you choose. Pay for description only when you need description.

    Can I extract a frame and immediately treat it as a new hero generate?

    You can attach it as a reference or as image-to-video frame zero. That is a second, priced job. The extract did not prepay it.

    What if every frame is bad?

    Then extraction did its job: it proved the take has no salvage. Reroll with a named change. Sampling the same failure at 12 fps will not invent a composition.

    Does this burn generation credits like a model?

    It is processing, not a catalog video model. Still not free of plan rules — and still the wrong button if what you wanted was a new render.