Generation controls

    Video input

    Also called Requires video, Footage-in, Video-in model.

    Video input is a catalog constraint: the model will not run unless you upload footage, as distinct from a still-in job or a text-only prompt.

    The live catalog splits on required inputs. A large class needs a photo. A smaller class needs a clip. A larger class needs only text. Seven models need both a still and footage. That cut is why /without exists: the missing input sends you down a different route set.

    Image-to-video takes a still and invents motion. Video-to-video restyles or edits footage you already shot. Motion-control drives a still with a driving clip — that is the both-required case. Treating a video-in model as a text-to-video row is how you get a refused job, not a cheaper one.

    The audio boolean on a row is not this flag. Native audio is what the output may contain. Video input is what you must bring.

    In practice

    • Read the required inputs on the model page before you write the prompt.
    • If you have a still and no clip, do not pick a video-in row; pick image-to-video or text-to-video.
    • Motion-control needs both a character still and a driving take — missing either is a different model.

    The mistake to avoid

    Pasting a prompt into a video-to-video or motion-control row and reading the refusal as a model failure. The ceiling was the missing file.

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.