Everything that would be a camera decision on a shoot is a wording decision here: focal length, distance, time of day, what fills the frame. Models differ enormously in how literally they take that wording — some treat the prompt as a shot list, others as a mood board — which is why the same sentence is not portable between them.
Stills are also where iteration is cheap. One frame renders in a fraction of the time a clip takes, so the sane pipeline for most video work is to settle composition, subject and palette as an image, then hand the finished frame to a video model.
The controls are different from video's. There is no duration and no motion, but there is frame shape, output size, and on many models a set of named style presets that shortcut a paragraph of prompt.
In practice
- Frame shape is set before generation — cropping afterwards throws away pixels you paid for.
- Typography and small text are the standard failure case; some models handle legible text, most do not.
- A locked seed plus a one-word prompt change is the cleanest way to see what a word is actually doing.
Text-to-image models
Catalog entries that render a still from a written prompt. 71 of the 296 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| GPT Image 2 Text to Image | OpenAI | Image |
| Nano Banana 2 | Image | |
| Seedream 5.0 Pro | ByteDance | Image |
| Grok Imagine Image Quality | Grok | Image |
| HiDream O1 Image | HiDream | Image |
| Luma UNI 1 Max | Luma | Image |
| Flux 2 Max | Flux | Image |
| Kling Image 3.0 | Kling | Image |
Browse all 30 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Stacking twenty adjectives. Past a point the model averages them, and the result is a picture that satisfies every word a little and none of them clearly.
Where you will run into it
- Text to Image Generator — One prompt. Every image model. One studio.
Related terms
Image-to-image
Image-to-image takes a picture as its primary input and returns a changed picture — restyled, corrected or varied — instead of inventing one from nothing.
Prompt
A prompt is the written instruction a generative model reads to decide what to make — the one input almost every model requires.
Aspect ratio
Aspect ratio is the proportion between an output's width and its height, written as two numbers — 16:9 for landscape, 9:16 for a full phone screen.
Seed
A seed is the number that decides the random starting noise for a generation, so the same seed with the same settings reproduces the same output.
Resolution
Resolution is how many pixels an output contains, usually named by its height — 720p, 1080p, 4K — and set before generation rather than after.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.