Industry

    Arena shopping and the '#1 model' claim

    Vendor number-one claims name whichever board flatters them. A two-minute check on the board, the date, the vote count and the exact category behind the claim.

    Versely Team7 min read

    In the first half of August 2026, two different companies could accurately describe their video model as the best available, and neither would be lying. Alibaba's Wan 3.0 sits first on Artificial Analysis's text-to-video board for models with audio, at 1243. Google's Gemini Omni Flash sits first on Arena's text-to-video leaderboard, at 1512. Wan 3.0 does not make Arena's top ten at all.

    This is not a scandal. It is what happens when two organisations measure with different vote pools, different model sets and different prompt distributions. But it does mean a marketing team can pick the board that produces the headline they want, and the reader has no way to tell whether that happened. The fix is a check that takes about two minutes and that almost nobody runs.

    What arena shopping looks like in practice

    The tell is a claim that names a rank without naming what it ranked. "The number one video model," "top-rated image model," "leading multimodal model." Every one of those is a rank stripped of the four things that would let you verify it.

    Here are five real placings from the same fortnight, each of which could support a very different press release.

    The claim you'd write What is actually true The part left out
    Number one video model Wan 3.0, first on Artificial Analysis's with-audio board at 1243 Not in Arena's top ten; public beta, full API described as coming
    Number one video model Gemini Omni Flash, first on Arena's text-to-video board at 1512 Second on Artificial Analysis's board, at 1238
    Number two at launch FLUX 3 video, second on Arena's text-to-video board at 1494 Debut placing built on roughly 1,300 votes
    Top five video model Meta's Muse Video, fifth on Arena's text-to-video board at 1457 Previewed on 7 July 2026 and still not publicly released
    Number one image editing model Reve 2.1, first on Artificial Analysis's image editing board at 1262 Second, not first, on that same organisation's text-to-image board at 1321

    The Reve row is the honest version of the problem. It is a genuinely strong model with two accurate headlines available, and choosing the flattering one is not fraud. It is just a claim that means less than it sounds like it means, and the reader deserves to know which board produced it.

    The two-minute check

    Four questions, in this order, each answerable from the board itself.

    1. Which organisation's board? (45 seconds.) Open the leaderboard the claim names and find the model. If the claim does not name an organisation, that is your answer already. Arena and Artificial Analysis are separate companies, and Artificial Analysis calling its boards "Image Arena" and "Video Arena" is the single largest source of misattribution in circulation.

    2. What is the date of the snapshot? (15 seconds.) Compare the claim's date to the board today. Both organisations add models continuously, and the video boards in particular took several new entrants across July and early August. A rank from six weeks ago has often been overtaken twice, and a claim carrying no date is one nobody can re-run.

    3. How many votes is it built on? (15 seconds.) Arena's text-to-video board carried 616,845 votes across 45 models as of 14 August 2026. A debut entry with roughly 1,300 votes sits on the same board as entries with orders of magnitude more, and the leaderboard renders both as a single number. A thin sample is not a dishonest sample, but it is a provisional one.

    4. Which exact category and configuration? (45 seconds.) Text-to-video, image-to-video, text-to-video-with-audio and video editing are separate boards, and models place very differently across them. Arena's entries also carry configuration in the name: dreamina-seedance-2.0-720p and dreamina-seedance-2.5-720p are separate rows at third and fourth, and veo-3.1-audio is the audio configuration rather than the model family. A claim that says "Seedance" or "Veo" has quietly merged rows.

    If all four answers check out, the claim is real. That is worth knowing too, and it happens more often than the cynical read suggests.

    The fifth question, which nobody asks

    Is it callable?

    Ranking and availability are completely independent, and the boards do not distinguish. Three examples from the current top ten alone:

    • Muse Video is fifth on Arena's text-to-video board. Meta previewed it on 7 July 2026 and access was described as coming soon to creators and in Meta AI. There is no product to buy.
    • Seedance 2.5 shipped on 31 July 2026 to Jimeng AI and Doubao Pro, which are China-first consumer surfaces. As of 2 August 2026 there was no callable Ark API endpoint, so outside China you cannot build on it regardless of where it ranks.
    • sora-2-pro is eighth at 1364. OpenAI announced Sora's discontinuation on 24 March 2026, the app closed on 26 April, and the API shuts on 24 September 2026. A top-ten rank and roughly five weeks of remaining API life are both true statements about the same entry.

    There is a related version of this for open-weights claims. Alibaba released Qwen-Image-3.0 on 21 July 2026 and made it generally available on 5 August with no weights, no licence, no technical report and no model card, a break from the Apache-2.0 releases that preceded it. If your plan depended on self-hosting, the rank was never the constraint.

    The question to add to any vendor conversation is therefore not "where does it rank" but "where does it rank, and can I call it from where I am, today."

    What the check still cannot tell you

    Nothing on a leaderboard measures your brief. A rank measures aggregate preference across a broad, deliberately generic prompt distribution, which is exactly the wrong shape for a question like "which model renders our packaging text legibly at 9:16."

    The replacement is small enough to actually build. Take five prompts you genuinely run, write down what a pass looks like for each, and run the shortlist against them. The agent chat will fan one prompt across several named models in a single request, which keeps the prompt fixed and the model as the only variable, and every model in the catalogue publishes its own specifications and credit cost so you can score quality against spend rather than against a headline. The compare pages are useful for narrowing a shortlist before you spend anything, and quality per credit is a better organising question than rank for most production work.

    Two practical notes on running it. Judge prompt adherence separately from aesthetics, because a model that looks better and follows instructions worse is a trap on any brief with brand constraints. And re-run the eval when you re-evaluate, not when a vendor publishes a claim. Your five prompts are stable; the boards are not.

    FAQ

    Are vendors lying when they claim number one?

    Usually not. In almost every case examined here the underlying placing is real and checkable. What gets omitted is which organisation's board, which category, and how thick the sample is, and those omissions are what turn an accurate placing into a misleading impression. Treat it as selective rather than false, and check rather than dismiss.

    Which board should I trust if they disagree?

    Neither, exclusively. When two boards rank the same category almost inversely, the honest conclusion is that the models are close enough that measurement choices dominate, which is precisely when your own eval decides it. Use the disagreement as a signal to test rather than as a puzzle to resolve.

    Does a higher Elo mean a better model for my work?

    Only in proportion to how much your work resembles the board's prompt distribution, which for most commercial briefs is not much. Elo aggregates strangers' preferences on generic prompts. It is a good shortlist generator and a poor decision-maker. Elo rankings explained covers what the number is actually counting.

    How often should I redo this?

    Quarterly for a production stack, and immediately when a model you depend on gets a deprecation notice. The Sora timeline is the argument for the second half of that: a model can hold a top-ten rank and have a hard API shutdown date at the same time, and the leaderboard will never tell you.