Generation modes

    Text-to-image meaning

    Also called T2I.

    Text-to-image renders a still picture from a written description, with no picture going in.

    Everything that would be a camera decision on a shoot is a wording decision here: focal length, distance, time of day, what fills the frame. Models differ enormously in how literally they take that wording — some treat the prompt as a shot list, others as a mood board — which is why the same sentence is not portable between them.

    Stills are also where iteration is cheap. One frame renders in a fraction of the time a clip takes, so the sane pipeline for most video work is to settle composition, subject and palette as an image, then hand the finished frame to a video model.

    The controls are different from video's. There is no duration and no motion, but there is frame shape, output size, and on many models a set of named style presets that shortcut a paragraph of prompt.

    In practice

    • Frame shape is set before generation — cropping afterwards throws away pixels you paid for.
    • Typography and small text are the standard failure case; some models handle legible text, most do not.
    • A locked seed plus a one-word prompt change is the cleanest way to see what a word is actually doing.

    Text-to-image models

    Catalog entries that render a still from a written prompt. 73 of the 331 models in the Versely catalog qualify.

    Browse all 44 spec pages for full settings, resolutions and credit costs.

    The mistake to avoid

    Stacking twenty adjectives. Past a point the model averages them, and the result is a picture that satisfies every word a little and none of them clearly.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.

    All terms →