Comparisons

    Wan 3.0 tops one video board, misses another

    Wan 3.0 leads Artificial Analysis' with-audio board and is absent from Arena's top ten. Why a missing model is a roster fact rather than a result.

    Versely Team7 min read

    Alibaba's Wan 3.0 is the number one text-to-video-with-audio model on Artificial Analysis, at 1243. It does not appear anywhere in the top ten of Arena's text-to-video board, which runs from 1512 down to 1363 across 45 models and 616,845 votes as of 14 August 2026.

    Both statements are true. Neither is a mistake. And the interesting part is not that two boards disagree, because the striking thing about these two boards is how much they agree. The disagreement is exactly one model wide.

    The two boards mostly agree, which is the surprising part

    Arena (arena.ai, rebranded from LMArena on 28 January 2026) and Artificial Analysis (artificialanalysis.ai) are unrelated companies running separate vote pools on separate rating scales, so 1243 on one and 1477 on the other are not comparable numbers. Take that as read, because something more interesting shows up once you line the two rosters against each other.

    Model Artificial Analysis (text-to-video with audio) Arena (text-to-video)
    Wan 3.0 1st, 1243 Not in the top ten
    Gemini Omni Flash 2nd, 1238 1st, 1512
    MiniMax H3 3rd, 1227 6th, 1453
    Seedance 2.0 720p 4th, 1220 3rd, 1482

    Three of the four land in Arena's top six in broadly the same order. Gemini Omni Flash is at or near the top of both. Seedance 2.0 is top four on both. MiniMax H3 is mid-pack-top on both. Two independent vote pools with different prompt distributions converged on nearly the same ordering, which is a reasonable sign that both are measuring something real.

    Wan 3.0 is the single outlier. And before reaching for "one board is wrong," there is a duller explanation sitting in plain sight.

    Absent from a board is not the same as beaten on it

    Artificial Analysis added Wan 3.0 to its board last month, alongside Vidu Q3 Turbo and MiniMax H3. Wan 3.0 entered public beta on 6 August 2026 and is reachable through Model Studio and the Wan site, with a full API described as "soon."

    A model that entered beta on 6 August, is not generally callable, and does not appear in the top ten of a 45-model board has a much more boring explanation available than "it lost": it may not be on that roster yet. Boards do not add models simultaneously, and a model that is hard to access programmatically is hard for a board to serve into blind matchups at volume.

    This matters because the two claims read identically in a headline and mean completely different things:

    • "Wan 3.0 ranked below the top ten on Arena" would be a result.
    • "Wan 3.0 does not appear in Arena's top ten" is what is actually verifiable, and it is consistent with the model simply not being listed.

    Do not upgrade the second into the first. Every one of these boards is a moving roster, and recency of addition is a confound before it is a finding.

    Three other ways a roster gap gets misread

    Wan 3.0 is one instance of a general problem. Absence has several causes and every one of them looks identical from outside.

    The model was added and has not accumulated votes. A new entry needs matchups before its rating stabilises. Until it has them it can sit low, or sit unlisted, depending on the board's minimum-vote threshold. Neither position is a judgment about the model.

    The entry exists under a name you did not search for. Boards list the specific build they tested, not the product. Arena carries dreamina-seedance-2.0-720p and veo-3.1-audio rather than "Seedance 2.0" and "Veo 3.1". A model you believe is missing may be sitting there at a resolution or an audio variant you did not think to look for.

    The board cannot serve it. Blind matchups at volume need programmatic access. A model reachable only through a gated beta console is expensive to include, which means access rather than quality decides whether it appears at all.

    None of those three is visible on the page. All three produce the same blank space, and a blank space is the one thing a leaderboard genuinely cannot explain about itself.

    The scoped-audio difference between the two boards is real and worth naming for completeness. Artificial Analysis' board is explicitly text-to-video with audio, so native audio is inside every matchup. Arena's headline board treats audio as a property of particular entries instead — veo-3.1-audio sits at ranks nine and ten as its own entry rather than folded into one audio-inclusive score. But that difference explains ordering. It cannot explain a model that is not on the list.

    And if your output is a person delivering a line to camera, neither board is answering your question in the first place. Both vote on text-to-video preference. A talking head is a lip-sync problem, and the models built for talking heads are frequently not the ones near the top of a general board.

    What this leaves you with on Wan

    Strip the ranking question out and Wan 3.0 reduces to a capability question with a much cleaner answer.

    Its headline is native 30-second clips to 1080p with audio in one pass, plus Omni-Reference, which accepts documents, spreadsheets, slides, PDFs and webpages as generation input. That second one is genuinely unusual and has no equivalent elsewhere in the current lineup. It is also in public beta without a general API, which means it is not something to build a pipeline on yet, whatever any board says about it.

    So the ranking was never the binding constraint here. Access was. An Elo rating can only narrow a list; it cannot tell you whether the thing at the top of that list is reachable, and in this case the top of one list is not.

    The callable end of the family today is Wan 2.7, which generates from text with motion consistency, optional custom audio input and prompt rewriting. If the audio ships with the picture in your brief, models with native audio is the shortlist worth starting from, and the callable options there include MiniMax H3 at 2K with native stereo and the Happy Horse family at 1080p with multilingual lip-sync. Test those against the actual brief while Wan 3.0's access story resolves.

    FAQ

    How do I tell whether a model is absent or just losing?

    Look for the entry, not the rank. If the name appears nowhere on the full board, that is a roster fact. If it appears low down, that is a result. Both boards publish full lists rather than just top tens, so checking takes about fifteen seconds — which is fifteen seconds more than most people spend before writing "ranked below."

    Why are the score ranges so different between the two boards?

    Because they are independently calibrated rating systems with different vote pools and different model sets. Artificial Analysis' published top four spans 1243 to 1220, a 23-point range. Arena's top ten spans 1512 to 1363. The scales are not interchangeable and converting between them is not meaningful.

    Is Wan 3.0 worth waiting for?

    Its distinctive feature, Omni-Reference over documents and structured files, has no substitute in the current callable lineup, so if that specific input path is what you need, it is worth tracking. If you need long single-pass generation with audio, other models reach that today and waiting costs you more than switching later would.

    Does the with-audio board mean those models have better audio than the Arena leaders?

    No. It means audio was part of what voters were grading in those matchups. It does not establish a ranking of audio quality on its own, and it does not tell you anything about models absent from that roster.