What you say to the agent
No special syntax — just describe it like you would to a person.
What it does, step by step
- 1
Give the agent the video and, optionally, the time range and how many frames you want.
- 2
It extracts the still(s) as image URLs.
- 3
Use the frame as a thumbnail, or chain it straight into a photo-edit or image-to-video request as a reference.
What it needs from you
- •A video to work from
What comes back
One or more still image URLs extracted from the video.
What it costs
A lightweight processing step, not a generative model — check your plan for details.
Under the hood
This is what the agent actually calls when you ask for it — real tools from its live surface, not marketing copy.
extract_video_framesExtract one or more still frames from a video as image URLs. Useful for grabbing a thumbnail, or chaining a frame into generate_image_from_image / generate_video_from_image as a reference/first frame.
See it done in a real workflow
Morning Matcha Routine
UGC-style wellness reel — a clean-girl creator walks through her actual morning matcha ritual, then cuts to a cozy animated hero shot of the finished iced latte.
TrendingNYC Street Interview
Photorealistic street-vlog interview — Riley works three different NYC corners at golden hour asking strangers one question: "What's the wildest thing you've ever done?" Three candid OTS/two-shot clips with locked character references, real handheld energy and native spoken dialogue. Vertical 9:16.
TrendingPrimo Protein vs Other Brand
Pixar-style 3D comparison ad — the confident PRIMO PROTEIN pouch faces off against the tired old rival brand's tub across 11 talking clips: pasture vs dusty pantry, herb garden vs toxic lab, clean American lab vs grimy factory, chocolate-milkshake CTA vs "wet sand." ~40s vertical reel with native character voices.
Or start from a one-tap template
Frequently asked questions
What do I actually say to the agent to extract frames from a video?+
Just describe it in plain English — for example: "Grab a still frame from this video at the 5-second mark for a thumbnail" The agent handles picking the right tool and model from there.
What does the agent need from me first?+
At minimum: A video to work from. Anything else it needs, it asks for before running.
What do I get back?+
One or more still image URLs extracted from the video.
Does this cost credits?+
A lightweight processing step, not a generative model — check your plan for details.
You can also just ask for
Ask your Versely agent to extract frames from a video
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.