Reading Image Arena Leaderboards Without Being Fooled
An Elo number tells you which image strangers preferred in blind votes, not which model will nail your specific job. Here's the gap, and how to close it.
Ask most creators why they picked an image model and the answer is some version of "it was ranked first." That's a reasonable instinct built on a fragile assumption: that a leaderboard rank measures the thing you actually care about. It doesn't, not directly. It measures something narrower and more specific, and the gap between what it measures and what you need it to mean is exactly where a lot of wasted generations come from.
What an Elo score is actually counting
Artificial Analysis's Image Arena produces its rankings from blind preference votes — a viewer sees two outputs from two different models, unlabeled, and picks the one they prefer. Do that across "millions of responses" and enough prompt variety, and you get an Elo rating: a single number per model, derived entirely from how often it won a head-to-head against whatever else it happened to be paired against.
That's a genuinely useful signal, and it's honest about what it is: an aggregate of strangers' snap preferences across a huge, varied, and — critically — generic set of prompts. It is not a measurement of how well a model does the one thing you need it to do repeatedly. Those are different questions, and a leaderboard only answers the first one.
Two specific ways that gap shows up matter enough to name.
Generation and editing are scored on different boards entirely
The most common mistake is treating "top-ranked image model" as one leaderboard when it's actually two. Artificial Analysis runs a separate leaderboard for image editing, distinct from the text-to-image arena, because editing an existing image — following an instruction while preserving everything you didn't ask to change — is a different skill than generating a new image from nothing. A model can rank near the top of one board and be unremarkable on the other, because nothing about generating a strong image from a blank canvas guarantees a model will hold your product's shape steady while it swaps the background.
If your actual job is "take this photo and change the lighting" and you're reading the text-to-image arena to pick a tool, you're consulting the wrong leaderboard — not a slightly-off one, a categorically different one measuring a different skill.
One number can hide the exact swing you'd want to see
The subtler problem shows up even when you're on the right board. Microsoft's MAI-Image-2.5 launched ranking No. 2 on Artificial Analysis's Image Edit leaderboard — a solid, legible headline result. But the model's own release notes report the real story sits underneath that rank: the largest category-level gains over its predecessor were +107 Elo in text rendering and +90 in cartoon/anime/fantasy output, against an overall average gain of +75 across all categories measured.
Sit with that gap for a second. The overall number — the one that produces the rank you'd actually see on the leaderboard — undersells what happened in text rendering by 32 points, and it has no way to tell you that if your job is specifically "get legible words inside a generated image," this release matters far more to you than the average creator scrolling the same board. A single rank is a weighted average across every prompt type voters happened to submit. If your work concentrates in one category — text-in-image, a particular art style, product photography, whatever it is — the aggregate score is diluting the exact signal you need with a hundred prompt types you don't run.
This isn't a knock on Elo as a method — a single comparable number across 150+ models is precisely what makes the arena useful for a first pass. It's a limit on what that number is for. It's built to answer "which model do voters prefer in general," and it answers that question well. It was never built to answer "which model is best at the specific thing I make," and treating it as though it does is the fooling part of "without being fooled."
Build the eval a leaderboard can't run for you
The fix isn't to distrust leaderboards — it's to use them for what they're actually good at (a fast first cut across dozens of models) and stop asking them to finish the job. The finishing move is a small private evaluation, and it's smaller than it sounds:
- Pull 6-10 prompts from your real work, not generic benchmarks. If you make product photography, use product photography prompts. If your job lives or dies on legible in-image text, every prompt should contain text.
- Shortlist candidates from the relevant board — generation or editing, whichever matches your actual task — rather than assuming rank transfers between them.
- Run the identical prompt set across your shortlist and judge the outputs against your own bar, not a stranger's blind preference. You're not looking for which one you like better in isolation; you're looking for which one is usably close to what you needed, most often.
- Weight the category you actually ship, not the average. A model that's mid-pack overall but strong specifically where your work concentrates will outperform the overall leader on your actual output, the same way MAI-Image-2.5's overall +75 undersold its +107 in the one category that might be all you care about.
Running the eval on Versely
Versely's model catalog and comparison tool are built for exactly this — checking a model's current rank by category before you commit a batch to it, rather than working from memory of whichever model was ahead last quarter. The part that actually saves time is that you don't need to run each candidate as a separate generation and eyeball them in different tabs: Versely's image generation tool accepts a list of models alongside one prompt, so the same request fans out across every model you name and returns them together.
A concrete version of the eval above, run directly in Versely's agent chat:
- Prompt: "Generate this exact product shot prompt across Recraft V4, Midjourney V7, and GPT Image 2 — I want to compare them side by side before I commit to a batch."
- The agent runs your prompt once across all three named models in a single request and returns the results together, rather than you re-prompting per model and losing the apples-to-apples comparison.
- Pick the winner for your prompt, not the leaderboard's average prompt, then run your actual production batch on it.
- If you want the category context first, check each candidate's current rank on the best AI image generator page or /compare before you shortlist — that's the fast first cut; the side-by-side generation is the finishing move.
That combination — public rank to shortlist fast, private side-by-side to decide for real — is the whole method. Neither step alone is enough: rank without your own eval picks the stranger-average winner, and an eval without a rank-based shortlist makes you test candidates at random instead of starting from the ones already known to be strong.
FAQ
Is Elo a bad way to rank AI image models?
No — it's a strong first-pass filter across a large field, precisely because it's a single comparable number derived from real blind preferences rather than a vendor's own marketing claims. The mistake is treating that one number as the final word on a specific job, when it's an average across every prompt type voters happened to submit.
Why do generation and editing leaderboards rank models differently?
Because they measure different skills. Generating a new image from text and editing an existing image while preserving what wasn't asked to change are different capabilities — a model strong at one has no guaranteed relationship to how it performs at the other, which is why Artificial Analysis scores them on separate boards.
How many prompts do I need for a useful private model eval?
Fewer than you'd think — 6 to 10 prompts drawn from your actual work is usually enough to see a clear pattern, as long as every prompt reflects the specific job you run repeatedly rather than generic test prompts. The goal is signal on your use case, not statistical rigor across every possible one.