Everything that would be a camera decision on a shoot is a wording decision here: focal length, distance, time of day, what fills the frame. Models differ enormously in how literally they take that wording — some treat the prompt as a shot list, others as a mood board — which is why the same sentence is not portable between them.
Stills are also where iteration is cheap. One frame renders in a fraction of the time a clip takes, so the sane pipeline for most video work is to settle composition, subject and palette as an image, then hand the finished frame to a video model.
The controls are different from video's. There is no duration and no motion, but there is frame shape, output size, and on many models a set of named style presets that shortcut a paragraph of prompt.
In practice
- Frame shape is set before generation — cropping afterwards throws away pixels you paid for.
- Typography and small text are the standard failure case; some models handle legible text, most do not.
- A locked seed plus a one-word prompt change is the cleanest way to see what a word is actually doing.
Text-to-image models
Catalog entries that render a still from a written prompt. 73 of the 331 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| GPT Image 2 Text to Image | OpenAI | Image |
| Grok Imagine Image 2.0 | Grok | Image |
| Nano Banana 2 | Image | |
| Seedream 5.0 Pro | ByteDance | Image |
| GPT Image 1.5 | OpenAI | Image |
| Grok Imagine Image Quality | Grok | Image |
| Qwen Image 3 Text to Image | Qwen | Image |
| Luma UNI 1 Max | Luma | Image |
Browse all 44 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Stacking twenty adjectives. Past a point the model averages them, and the result is a picture that satisfies every word a little and none of them clearly.
Where you will run into it
- Text to Image Generator — One prompt. Catalog image models. One studio.
Related terms
Image-to-image
Image-to-image takes a picture as its primary input and returns a changed picture — restyled, corrected or varied — instead of inventing one from nothing.
Prompt
Prompt meaning: the written instruction a generative model reads to decide what to make. The one input almost every model needs.
Aspect ratio
Aspect ratio meaning: width-to-height proportion of an output, written as two numbers (16:9 landscape, 9:16 full phone).
Seed
Seed meaning: the number that sets the random starting noise so the same seed and settings can reproduce the same output.
Resolution
Resolution is how many pixels an output contains, usually named by its height — 720p, 1080p, 4K — and set before generation rather than after.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free account. Works in your browser - no install needed. The same account signs in on your phone.