AI News

    LMArena is now Arena: fix your citations

    LMArena became Arena in January, and Artificial Analysis is a separate company. The five-field citation rule that stops a rank going to the wrong board.

    Versely Team7 min read

    On 28 January 2026, LMArena renamed itself Arena and moved to arena.ai. Seven months later, "LMArena" is still what a great many briefs, client decks and vendor blog posts say, and some of those attribute a number that did not come from Arena at all.

    That second failure is the expensive one. Artificial Analysis is a separate company running its own boards, including an Image Arena and a Video Arena. The shared word does most of the damage: someone reads "Video Arena," writes "the arena," and the claim ends up credited to whichever organisation the reader assumes. Right now those two organisations disagree about which video model is best, so the assumption decides the answer.

    Two organisations, one overloaded word

    Here is the distinction in the smallest form that stays correct.

    Arena Artificial Analysis
    Domain arena.ai artificialanalysis.ai
    Former name LMArena, until 28 January 2026 No rename
    Video board cited here Text-to-video leaderboard Text-to-video leaderboard for models with audio
    Relationship Unrelated companies Unrelated companies

    Artificial Analysis calls its boards "Image Arena" and "Video Arena." Arena is a company name. Those two facts collide in every casual citation, and there is no way to be precise about it except to write the organisation's name out.

    The numbers are not on the same scale

    This is the part that turns a naming slip into a wrong statement. Both organisations publish Elo ratings, which are computed from blind head-to-head preference votes. But an Elo rating is only meaningful relative to the pool it was computed against, and these are different pools, different model sets and different prompt distributions.

    As of 14 August 2026, Arena's text-to-video leaderboard carried 616,845 votes across 45 models, with gemini-omni-flash first at 1512 and flux-3-video second at 1494. On the same date, Artificial Analysis's text-to-video board for models with audio had Wan 3.0 first at 1243, Gemini Omni Flash second at 1238 and MiniMax H3 third at 1227.

    A 1512 and a 1243 are not comparable. They are not on the same axis, they are not calibrated against each other, and the difference between them says nothing whatsoever about model quality. If you have ever seen a comparison table that lists scores from both organisations in the same column, that table is meaningless, and the person who built it almost certainly did not realise.

    The five-field citation rule

    A leaderboard claim is only checkable if it carries five things. Drop any of them and the reader cannot reconstruct what you saw.

    1. The organisation. "Arena" or "Artificial Analysis," spelled out. Never "the arena," never "LMArena," never "the leaderboard."
    2. The board. Text-to-video is not image-to-video is not text-to-video-with-audio is not image editing. Models place very differently across them, which is the entire reason the boards are separate.
    3. The date. Boards move weekly as new models are added and votes accumulate. Artificial Analysis added Wan 3.0, Vidu Q3 Turbo and MiniMax H3 to its video board within the last month, and Ideogram P-Image variants, Qwen-Image-3.0 and MAI-Image-2.5-Pro to its image board. A rank without a date has a shelf life you cannot see.
    4. The sample. Vote count for the board, and ideally for the entry. A placing built on a few thousand votes has a much wider confidence interval than one built on hundreds of thousands.
    5. The model string as the board writes it. Arena's video board lists dreamina-seedance-2.0-720p and dreamina-seedance-2.5-720p as separate entries, and veo-3.1-audio rather than "Veo 3.1." Those suffixes are configuration. Dropping them changes the claim.

    Applied, the difference looks like this.

    Weak: "Seedance is the number three video model on LMArena."

    Four problems. The organisation is called Arena. "Seedance" is ambiguous between two entries that sit at third and fourth. "Video" does not identify the board. And there is no date, so the claim quietly ages into being false.

    Checkable: "As of 14 August 2026, dreamina-seedance-2.0-720p ranked third on Arena's text-to-video leaderboard at 1482, on a board carrying 616,845 votes across 45 models."

    That sentence is longer. It is also auditable by anyone who wants to, which is the standard a claim in a client deck should meet.

    Two traps the rule catches

    The debut placing. Black Forest Labs' FLUX 3 entered Arena's text-to-video board at second place on roughly 1,300 votes. That is a real result and a genuinely strong debut, but a placing built on 1,300 votes on a board with 616,845 total is a different quality of evidence from a placing built over months. The fourth field, the sample, is what surfaces that difference. If you want to try the model rather than argue about its rank, Flux 3 Text to Video and Minimax H3 Text to Video are both live in the catalogue.

    The anonymous entry. Alibaba's HappyHorse 1.0 debuted anonymously on Artificial Analysis around 7 April 2026 and reached first place before the lab behind it was identified. Any citation written in that window attributed a top rank to a developer nobody could name. Worth a sixth question when a codename is involved: was the entry's developer known on the date of the snapshot?

    Where a wrong attribution actually costs you

    Three places, in rough order of how much it hurts.

    Ad claims. A superiority claim in paid creative that names the wrong source is a substantiation problem, not a typo. The number may be real and still fail because it did not come from where you said it did.

    Client decks. The specific failure is a client who checks. They search the organisation you named, find a different number on that organisation's board, and now every other figure in the deck is suspect.

    Internal model selection. This is the quiet one. A team standardises on a model because it was "first," and the board it was first on measured something other than the job. Gemini Omni Flash leads one organisation's video board and sits second on the other's, which is a perfectly good reason to shortlist Gemini Omni Video and a bad reason to skip testing it against your own brief.

    The general fix is to treat a leaderboard as a shortlist generator rather than a decision. Reading image arena leaderboards without being fooled goes through what an Elo does and does not measure for your specific job, and running the shortlist as a head-to-head on your own prompts is what the compare pages and the text-to-video shortlist are for.

    FAQ

    Is Arena the same organisation as Artificial Analysis?

    No. They are unrelated companies. Arena, at arena.ai, was called LMArena until 28 January 2026. Artificial Analysis, at artificialanalysis.ai, runs its own boards and happens to call two of them Image Arena and Video Arena, which is where most of the confusion originates.

    Can I compare an Elo rating from one to an Elo rating from the other?

    No. Each rating is computed within its own vote pool against its own set of models. The absolute numbers have no shared calibration, so a higher figure on one board does not mean a better model than a lower figure on the other. Compare ranks within a single board, and only against models that appear on that same board.

    What should I write if I cannot find the vote count?

    Give the four fields you have and say the sample is unstated. "As of 14 August 2026, first on Arena's text-to-video leaderboard; board-level vote count not recorded at time of citation" is honest and still checkable. Omitting the field silently is what makes a thin result look like a thick one.

    Do I need to update citations I already published?

    Update the ones that are still doing work: pitch decks in circulation, landing pages, anything with a superiority claim attached. The minimum fix is renaming LMArena to Arena and adding the snapshot date. If a claim cannot be re-verified on the board it named, the right move is to pull the number rather than restate it.