The live catalog splits on required inputs. A large class needs a photo. A smaller class needs a clip. A larger class needs only text. Seven models need both a still and footage. That cut is why /without exists: the missing input sends you down a different route set.
Image-to-video takes a still and invents motion. Video-to-video restyles or edits footage you already shot. Motion-control drives a still with a driving clip — that is the both-required case. Treating a video-in model as a text-to-video row is how you get a refused job, not a cheaper one.
The audio boolean on a row is not this flag. Native audio is what the output may contain. Video input is what you must bring.
In practice
- Read the required inputs on the model page before you write the prompt.
- If you have a still and no clip, do not pick a video-in row; pick image-to-video or text-to-video.
- Motion-control needs both a character still and a driving take — missing either is a different model.
The mistake to avoid
Pasting a prompt into a video-to-video or motion-control row and reading the refusal as a model failure. The ceiling was the missing file.
Related terms
Image-to-video
Image-to-video animates a still you supply: the picture becomes the opening frame, and the prompt describes only what happens next.
Video-to-video
Video-to-video takes finished footage in and returns altered footage, using the original clip as the structural reference for every frame.
Motion control
Motion control transfers the movement in a driving video onto a different subject — the performance stays, the performer changes.
Reference-to-video
Reference-to-video builds a clip around subjects supplied as separate reference images, rather than starting from one fixed opening frame.
Seed
A seed is the number that decides the random starting noise for a generation, so the same seed with the same settings reproduces the same output.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.