Generation modes

    Text-to-video

    Also called T2V.

    Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.

    With nothing to copy, the model decides everything: the subject, the set, the lens behaviour, the light. That makes a text-to-video prompt do two jobs at once — it picks what appears and it picks how the thing is shot. A prompt that only names a subject has quietly handed the second job back to the model, which is why three words return a different film every time you press go.

    The trade is control. You cannot guarantee that a face, a product label or a brand colour survives from one generation to the next, because there is no reference for it to survive from. Iteration is the tool instead: hold the seed, change one clause, and compare the two takes side by side.

    It wins on establishing shots, abstract b-roll and anything where a version of the idea is good enough. It loses the moment the clip has to match something that already exists — a real product, a real face, a shot you generated yesterday. Those are jobs for image-to-video and reference-to-video.

    In practice

    • The prompt carries subject, action, camera behaviour, light and the state the shot ends in.
    • Clip length is chosen up front from a fixed set of options, not trimmed afterwards.
    • The same prompt at a different seed is a different take, not a correction.

    Text-to-video models

    Catalog entries that will build a clip from a written prompt alone. 50 of the 296 models in the Versely catalog qualify.

    ModelProviderType
    Happy Horse 1.0 Text to VideoAlibabaVideo
    Seedance 2.0ByteDanceVideo
    Wan 2.7 Text to VideoWanVideo
    Kling 2.5 TurboKlingVideo
    Grok Imagine VideoGrokVideo
    Vidu Q3 VideoViduVideo
    Runway Gen-4.5RunwayVideo
    VEO 3.1GoogleVideo

    Browse all 32 spec pages for full settings, resolutions and credit costs.

    The mistake to avoid

    Reading a weak result as a model failure. Far more often the prompt named a subject and nothing about the shot, so framing, motion and lighting were all left to chance.

    Go deeper

    Text-to-Video: A Beginner's Guide to AI Video Generation

    Everything a new creator needs to know about text-to-video AI in 2026 — how the models work, which one to pick, prompt patterns that actually generate usable clips, and the pitfalls to avoid.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.