It comes from chess, and the borrowed logic is the point. Beating a highly-rated opponent moves your number more than beating a weak one, so a rating reflects the quality of the field a model has faced rather than a raw win count. It is a relative measure — it only means something against the population it was computed over.
It exists because the alternatives are worse. Automated image metrics correlate poorly with what people actually prefer, and a single expert's opinion does not generalise. Aggregated blind pairwise votes are noisy per vote and reasonably stable in bulk.
The caveats are real. Ratings drift as new models enter, prompt mix changes what is being measured, and preference rewards what looks striking in a side-by-side — which is not always what survives in a finished edit. A ranking is a starting shortlist, not a verdict on your brief.
In practice
- Relative, not absolute: a number only means something against the same field.
- Small gaps between neighbouring models are usually inside the noise.
- Category rankings beat overall rankings when your job is specific.
Models carrying a ranking
Catalog entries with at least one placed result in the live rankings. 148 of the 296 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| GPT Image 2 Text to Image | OpenAI | Image |
| Mai Image 2.5 Edit | Microsoft | Image |
| Happy Horse 1.0 Text to Video | Alibaba | Video |
| Gemini 3.1 Flash TTS | Audio | |
| Nano Banana 2 | Image | |
| Seedance 2.0 | ByteDance | Video |
| Cartesia Sonic 3.5 | Cartesia | Audio |
| Seedream 5.0 Pro | ByteDance | Image |
Browse all 77 spec pages for full settings, resolutions and credit costs.
The mistake to avoid
Picking the top-rated model for every job. Ratings aggregate across prompt types, and the leader on general preference is often not the leader on your particular one.
Go deeper
ELO Rankings for AI Models, Explained
How ELO rankings for AI models work: pairwise voting, rating math, what leaderboard gaps mean, and the blind spots to know before trusting a score.
Related terms
Prompt adherence
Prompt adherence is how faithfully a model does what the prompt actually said, as opposed to producing something attractive in the same neighbourhood.
Distillation
Distillation trains a smaller or faster model to imitate a larger one's outputs, which is where the fast and turbo variants of familiar models come from.
Multimodal model
A multimodal model handles more than one kind of data — text, images, audio, video — inside a single system rather than bolting separate tools together.
Diffusion model
A diffusion model generates by starting from random noise and removing a little of it at a time until a picture or clip is left behind.
Diffusion transformer
A diffusion transformer is a diffusion model whose internals are a transformer — the same architecture behind large language models — instead of the convolutional network earlier image models used.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.