Generation modes

    Text-to-image

    Also called T2I.

    Text-to-image renders a still picture from a written description, with no picture going in.

    Everything that would be a camera decision on a shoot is a wording decision here: focal length, distance, time of day, what fills the frame. Models differ enormously in how literally they take that wording — some treat the prompt as a shot list, others as a mood board — which is why the same sentence is not portable between them.

    Stills are also where iteration is cheap. One frame renders in a fraction of the time a clip takes, so the sane pipeline for most video work is to settle composition, subject and palette as an image, then hand the finished frame to a video model.

    The controls are different from video's. There is no duration and no motion, but there is frame shape, output size, and on many models a set of named style presets that shortcut a paragraph of prompt.

    In practice

    • Frame shape is set before generation — cropping afterwards throws away pixels you paid for.
    • Typography and small text are the standard failure case; some models handle legible text, most do not.
    • A locked seed plus a one-word prompt change is the cleanest way to see what a word is actually doing.

    Text-to-image models

    Catalog entries that render a still from a written prompt. 71 of the 296 models in the Versely catalog qualify.

    ModelProviderType
    GPT Image 2 Text to ImageOpenAIImage
    Nano Banana 2GoogleImage
    Seedream 5.0 ProByteDanceImage
    Grok Imagine Image QualityGrokImage
    HiDream O1 ImageHiDreamImage
    Luma UNI 1 MaxLumaImage
    Flux 2 MaxFluxImage
    Kling Image 3.0KlingImage

    Browse all 30 spec pages for full settings, resolutions and credit costs.

    The mistake to avoid

    Stacking twenty adjectives. Past a point the model averages them, and the result is a picture that satisfies every word a little and none of them clearly.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.