Reve 2.1 leads the image editing board
Reve 2.1 holds first on Artificial Analysis image editing at 1262 Elo. What an editing score actually measures, and how to turn a rank into a revision workflow.
Reve 2.1 sits first on Artificial Analysis' image editing leaderboard at 1262 Elo. GPT Image 2 (high) is second at 1258, and Microsoft's MAI-Image-2.5 is third at 1254. Four points between first and second; eight between first and third.
That spread is the most useful thing on the board, and not for the reason it looks like. Four points on a preference-vote scale is not a capability difference you will perceive on your own work. It is a coin flip that landed. Reading "number one on image editing" as a procurement decision is how teams end up switching editing models every six weeks and rebuilding their instruction library every time.
What is worth extracting from that board is narrower and more durable: a shortlist, and a clear understanding of what the score is scoring. Both of those survive the next reshuffle.
What an editing Elo is actually scoring
A text-to-image score comes from a simple comparison: two outputs, one prompt, pick the one you prefer. An editing score comes from a harder one. The voter sees a source image, an instruction, and two edited results, and picks a winner. That single vote silently bundles three separate judgments:
- Did the model do the thing that was asked? Instruction compliance. The narrow, obvious criterion.
- Did it leave everything else alone? Preservation. Change the sky and the product label should be untouched, the crop unchanged, the skin tone identical.
- Is the result attractive on its own terms? Aesthetic quality, entirely independent of the instruction.
Those three can trade against each other, and the vote gives you no way to see which one drove the result. A model that follows instructions brilliantly while quietly resampling the whole frame — shifting colour, softening detail, nudging a crop — can beat a model that changed exactly the requested region and nothing else, because the second image looks less impressive at a glance even though it is the correct edit.
That is the specific gap between an editing leaderboard and a production retouch workflow. Preference voters are grading a picture. You are grading a diff. We have written about reading arena leaderboards without being fooled in general terms; the editing board has this extra wrinkle on top, and it is the one that costs real money because a model that silently redraws unasked-for regions fails a brand audit that a preference voter would never run.
Why the top band is a tie, and what to decide on instead
Take 1262, 1258 and 1254 seriously as a band, not as an ordering. All three are in the top group. So, for practical purposes, is anything within roughly the same range. Elo is derived from win rates in head-to-head pairings, which means small differences reflect sampling as much as skill, and the gap widens or closes as votes accumulate.
Once you accept the top band is a tie, the tiebreakers become the things a leaderboard never measures:
- Preservation behaviour on your specific asset class. Product photography, faces, packaging with legal copy, and illustrated work each stress preservation differently.
- How it handles an instruction it cannot satisfy. Does it refuse, do nothing, or do something adjacent and wrong? The third is the expensive failure because it passes a quick glance.
- Whether it supports the control you need. Masking, reference images, batch application.
- Credit cost per revision round. Revision counts are the real driver of an editing budget, not the cost of any single edit. Every model page in the catalog shows its credit cost next to its rank.
- Consistency across a batch. A model that produces one excellent edit and forty inconsistent ones is worse than a model that produces forty identical adequate ones.
None of those appear on any board, and every one of them decides whether a model survives contact with a client revision cycle.
Turning a ranking into a revision workflow
Here is the sequence that converts a leaderboard into something operational. It takes about an hour once.
- Shortlist from the editing board, not the generation board. These are separate leaderboards measuring separate skills. Reve 2.1 is first on editing at 1262 and second on Artificial Analysis' text-to-image board at 1321, behind GPT Image 2 at 1369 — two different positions on two different boards for the same model. Take three to four candidates from the editing board and stop.
- Build a six-instruction suite from real revision notes. Not benchmark prompts. Pull the actual sentences clients have sent you: "make the background warmer," "remove the person on the left," "change the shirt to navy," "the logo needs to be bigger." Those are the instructions your model will run for the next year.
- Run all six across all candidates on the same source images. The comparison is only meaningful with the source held constant. The Versely agent can fan one instruction across several named models in a single request, which is what makes running the suite an hour rather than an afternoon.
- Score preservation separately from compliance. Two columns, not one. For each output, ask "did it do the thing" and "did it break anything else" as independent questions. This is the column a leaderboard cannot give you and the one that predicts client friction.
- Lock a winner per instruction class, not globally. Object removal, colour change, background replacement and text edits are different tasks. It is normal for two different models to win two of them, and treating your model choice as one decision instead of four leaves quality on the table.
- Write the winner into a documented instruction template. The compounding asset is the phrasing that works, not the model that ran it. When the board reshuffles — and it will — a documented template ports in an afternoon.
Preservation is the metric you have to build yourself
The single highest-value thing to add to that workflow is a preservation check, because nothing public measures it and it is where edited assets fail brand review.
The cheap version: for each candidate edit, look at three regions you did not ask to change. A logo, a face, and any fixed legal or numeric copy. If any of the three moved, softened, recoloured or subtly redrew, the edit failed regardless of how good the requested change looks. Barcodes and legal copy are the harshest test and the most consequential — what happens to barcodes and legal copy on generated packaging is worth reading before you trust any editing model with a package face.
The systematic version is the same check formalised into a written list against your brand manual, so a second person can run it without your judgment in the room.
Running it in the Versely catalog
Reve 2.1 itself is not in the Versely catalog. What is there is a set of editing models with pages you can read side by side, including MAI-Image-2.5 Edit and Nano Banana 2, which the catalog places third and fifth on that same editing board. Best image editing model is the ranked view across the whole set, and MAI-Image-2.5 Edit versus Nano Banana 2 is the head-to-head if your shortlist has already narrowed to two.
For actually running the suite, edit a photo with AI is the agent path. Run it in batch rather than one asset at a time — consistency across forty edits, not peak quality on one, is the metric that decides whether a model survives a real revision cycle.
FAQ
Is a four-point Elo gap meaningful?
Not for a purchasing decision. Elo differences of that size sit inside normal sampling variation and can invert as votes accumulate. Treat the models within a few points of each other as a band, shortlist all of them, and decide between them on preservation behaviour, control surface and credit cost.
Why can a model rank first on editing and not on text-to-image?
Because they are different skills scored on different boards. Generating a strong image from nothing is unrelated to following an instruction while leaving the rest of an existing image untouched. Reve 2.1 leading editing at 1262 while sitting second on text-to-image at 1321 is the normal shape of this, not an anomaly.
What does an editing leaderboard fail to capture?
Preservation of regions you did not ask to change, behaviour on instructions the model cannot satisfy, consistency across a batch, and credit cost across a multi-round revision cycle. Voters grade a single finished picture; production grades the difference between two pictures across many assets.
How often should I re-run the shortlist?
Once a quarter is enough for most teams, or whenever a model you actually use is deprecated. Chasing weekly board movement costs more in re-tuned instructions than the ranking difference is worth, which is exactly why the durable asset is the documented instruction template rather than the model name.