It comes from chess, and the borrowed logic is the point. Beating a highly-rated opponent moves your number more than beating a weak one, so a rating reflects the quality of the field a model has faced rather than a raw win count. It is a relative measure — it only means something against the population it was computed over.
It exists because the alternatives are worse. Automated image metrics correlate poorly with what people actually prefer, and a single expert's opinion does not generalise. Aggregated blind pairwise votes are noisy per vote and reasonably stable in bulk.
The caveats are real. Ratings drift as new models enter, prompt mix changes what is being measured, and preference rewards what looks striking in a side-by-side — which is not always what survives in a finished edit. A ranking is a starting shortlist, not a verdict on your brief.
In practice
- Relative, not absolute: a number only means something against the same field.
- Small gaps between neighbouring models are usually inside the noise.
- Category rankings beat overall rankings when your job is specific.
Models carrying a ranking
Catalog entries with at least one placed result in the live rankings. 181 of the 331 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| GPT Image 2.5 Sunburst Edit | OpenAI | Image |
| Cartesia Sonic 3.6 | Cartesia | Audio |
| Wan 3.0 Text to Video | Wan | Video |
| Gemini 3.8 Flash TTS | Audio | |
| GPT Image 2.5 Flare Edit | OpenAI | Image |
| Minimax H3 Text to Video | MiniMax | Video |
| GPT Image 2 Text to Image | OpenAI | Image |
| Grok Imagine Image 2.0 | Grok | Image |
Browse all 108 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Picking the top-rated model for every job. Ratings aggregate across prompt types, and the leader on general preference is often not the leader on your particular one.
Go deeper
ELO Rankings for AI Models, Explained
How ELO rankings for AI models work: pairwise voting, rating math, what leaderboard gaps mean, and the blind spots to know before trusting a score.
Related terms
Prompt adherence
Prompt adherence meaning: how faithfully a model does what the prompt said, not just something attractive nearby.
Distillation
Distillation trains a smaller or faster model to imitate a larger one's outputs, which is where the fast and turbo variants of familiar models come from.
Multimodal model
A multimodal model handles more than one kind of data — text, images, audio, video — inside a single system rather than bolting separate tools together.
Diffusion model
A diffusion model generates by starting from random noise and removing a little of it at a time until a picture or clip is left behind.
Diffusion transformer
A diffusion transformer is a diffusion model whose internals are a transformer — the same architecture behind large language models — instead of the convolutional network earlier image models used.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free account. Works in your browser - no install needed. The same account signs in on your phone.