Starting point · Without a camera

    How to Make Videos Without a Camera

    If you have nothing to upload, you are in the text-only lane: the prompt IS the footage. Pick a route by how long the finished video needs to be.

    no cameranothing to film withno footage to work withcan't film anything myselfno b-rollmake a video from just an idea

    The situation

    This page is for the specific situation where the shoot is not the bottleneck — there is no shoot. No camera, no phone footage worth using, no stock library, nothing on a drive to cut. That rules out every route that starts with "upload your clip", which is most published advice about making video with AI.

    The catalog splits cleanly on this. Of the 296 models Versely runs, 113 refuse to start without a still image and 22 refuse to start without footage. The 168 that require neither are the entire set available to you, and they are the only ones the routes below use.

    The three routes differ by output length, not by quality. A single clip, a multi-scene piece, and a narrated piece assembled around a voice track need different starting points, and picking the wrong one is the usual reason people conclude "AI video doesn't work" after one attempt.

    What the model catalog says about it

    Every model declares the input it requires, so “what you don’t have” is a filter over the catalog rather than a mood. These are counted from the live snapshot at build time.

    Models you can use

    168

    require no upload of any kind — a written prompt is the only input they need

    Models you can't

    135

    declare a required image or video input, so they never start from a blank page

    Published scene recipes

    166

    across 17 workflows — every image prompt and video prompt, already written

    Which route fits you

    Find the line that matches your situation, then read that route.

    The routes that work

    1

    Write the shot, not the shot list

    Use it when: You need one clip — a hook, an ad cutaway, a single moment — and you need it in one pass.

    1. 01Describe one continuous shot: subject, what it does, camera move, lighting, and the look. One shot, not a sequence — text-to-video models cut badly when a prompt asks for two locations.
    2. 02Set the aspect ratio before you generate, not after. Vertical output that was framed for 16:9 loses the subject's head.
    3. 03Generate two or three takes of the same prompt before you change the prompt. Variance between takes on one prompt is usually wider than the difference between two prompts.
    4. 04Keep the take that has the right motion, even if the details are wrong — details can be fixed downstream, motion can't.

    Run it here

    2

    Start from a premise and let it break into scenes

    Use it when: The finished thing is 30 seconds or longer and needs a beginning, a middle and an end.

    1. 01Write the premise as one sentence — who, what changes, and how it ends. Scene splitting works from a premise; it does not work from a mood.
    2. 02Lock the recurring characters, props and locations as named assets FIRST. Continuity across scenes comes from reused reference assets, not from repeating adjectives in every prompt.
    3. 03Generate scene by scene and accept that some scenes need four takes. A 6-scene piece where one scene is wrong reads as broken; the other five carry nothing.
    4. 04Stitch, then add the voice and captions over the finished cut rather than baking narration into each clip.

    Run it here

    3

    Cover a voice you already have

    Use it when: You can talk but you can't film — a recorded take, a podcast clip, or a script you're happy to read aloud.

    1. 01Cut the audio first and treat it as the spine. Everything visual is timed to it, so re-timing later means regenerating clips.
    2. 02Break the audio into beats and write one generated visual per beat — literal illustration of the sentence being spoken beats abstract footage on retention.
    3. 03Generate the visuals as separate short clips rather than one long one, so a bad beat costs one clip.
    4. 04Lay the clips over the audio, then caption. Captions last because the cut length is what decides where lines break.

    Run it here

    What this does not fix

    Routing around a missing input is not the same as not needing it. Three things stay broken.

    • It will not show your actual product. A generated object resembles the category, not your unit — if the video needs the thing you sell, read the product-sample route instead.
    • It will not show a specific real place. Text-only models render a plausible version of "a Brooklyn corner store", never yours.
    • It will not put you on screen. Generated people are not you, and pretending otherwise in a founder video is a trust problem, not a production one.

    Frequently asked questions

    Can you really make a video with no footage at all?+

    Yes — that is what the text-only class of models is. 168 of the 296 models in Versely's catalog declare no required image or video input, so a written prompt is the only thing they need to produce a clip. The constraint is length: a single generation is a few seconds, so anything longer is several generations stitched together.

    What's the difference between this and stock footage?+

    Stock gives you a real clip that was shot for somebody else, so you edit around what it happens to contain. A generated clip is written to your beat, so the subject, framing and action are what your script asked for — but it is synthetic, and it should not be presented as documentary footage of a real event.

    Which route should I start with if I've never made one?+

    The single-clip route. It has one failure mode you can see in 30 seconds, which makes it the fastest way to learn how much detail a prompt needs. Multi-scene work compounds prompt mistakes across every scene, so it is a bad first attempt.

    Do I still need a camera later?+

    Only for things that must be genuinely yours: your face, your premises, your physical product in a real hand. Everything that is illustrative — B-roll, transitions, hooks, establishing shots — has no honest reason to be filmed.

    If that wasn’t your constraint

    People usually arrive here missing one thing and leave realising it was a different one.

    Publishing on a schedule with the same gap?

    A route gets one video made. A format gets the next twenty made — same constraint, decided once.