Comparisons

    Which video benchmarks actually score audio

    Artificial Analysis runs a text-to-video board scoped to output with audio. Arena's board is not scoped that way, and the two rank almost inversely.

    Versely Team8 min read

    Two public leaderboards ranked text-to-video in mid-August 2026, and they disagreed so completely that the model sitting at No. 1 on one of them does not appear in the other's top ten. That is not noise in the Elo. It is two boards measuring different outputs and both calling the result "text-to-video."

    The difference that explains most of the gap is audio. One board is scoped to video with sound. The other is not. If your brief says sound-on, only one of them is answering your question, and the wrong citation will point you at a model that ranks well on silent footage.

    The two boards are run by different companies

    Start with the naming, because this is where most write-ups go wrong. LMArena rebranded to Arena on 28 January 2026 and now lives at arena.ai. Artificial Analysis is an unrelated company at artificialanalysis.ai running its own image and video arenas. Both use blind head-to-head preference voting. Both publish Elo ratings. They are not the same organisation, and their video boards do not agree.

    As of 14 August 2026, Arena's text-to-video board carried 616,845 votes across 45 models:

    Rank Model Elo
    1 gemini-omni-flash 1512
    2 flux-3-video 1494
    3 dreamina-seedance-2.0-720p 1482
    4 dreamina-seedance-2.5-720p 1477
    5 muse-video 1457
    6 minimax-h3 1453
    7 happyhorse-1.0 1428
    8 sora-2-pro 1364
    9–10 veo-3.1-audio 1364 / 1363

    Artificial Analysis publishes its text-to-video leaderboard explicitly scoped to output with audio. Its top entries on the same date: Wan 3.0 at 1243, Gemini Omni Flash at 1238, MiniMax H3 at 1227, Seedance 2.0 720p at 1220. Wan 3.0, Vidu Q3 Turbo and MiniMax H3 were all added to that board in the preceding month.

    Notice what Wan 3.0 does across the two. It is first on the audio-scoped board and absent from the other board's top ten entirely.

    Why the scoping changes the order

    On the Arena table, audio shows up as a naming convention rather than a scored dimension: veo-3.1-audio sits there as its own entry, distinct from the rest of the Veo line. That is what it looks like when audio is a model variant a voter may or may not have been comparing, rather than a property the board is built to measure. A voter picking between two clips on a board that is not scoped to sound can plausibly be judging motion, prompt adherence and artefacts, and nothing about whether the footsteps landed on the frames where feet hit the ground.

    An audio-scoped board pushes exactly the opposite way. It rewards models that were trained to emit picture and sound in one pass, and it flattens the advantage of a model whose strength is beautiful silent motion. That is not a flaw in either board. It is the boards doing what they say they do.

    The practical consequence: a model's rank moves substantially depending on whether the board counts the audio track, and every model with native audio is affected in the same direction. When a vendor blog claims a "#1 model" in July or August 2026 without naming the board, assume board-shopping until proven otherwise.

    Which board to cite for which brief

    Your brief Cite Because
    Sound-on social, dialogue, sync-critical Artificial Analysis text-to-video (with audio) Audio is inside the scored comparison
    Silent B-roll, footage you will score in post Arena text-to-video Sound is not the deciding variable
    A single "best model" claim in a deck Neither, alone Cite both and show the disagreement
    A model you can actually call today Neither See the availability check below

    That last row matters more than it looks. Two of Arena's top ten are not usable products for most teams. Muse Video sits at No. 5 and has never shipped — Meta previewed it on 7 July 2026 alongside Muse Image and has still only demoed it. sora-2-pro sits at No. 8 and is on a countdown: OpenAI announced Sora's discontinuation on 24 March 2026, the app shut on 26 April, and the API shuts on 24 September 2026. A leaderboard rank is not an availability signal.

    The check that beats both boards

    Neither arena knows what your footage has to do. The fix is the same one that applies to image leaderboards: use the board for the first cut, then run your own comparison on your own prompts.

    For sound-on work, a usable private eval is four prompts, not forty:

    1. Sync test. Something with a hard visual transient — a door closing, a bottle set on a counter, a single clap. Watch whether the sound lands on the frame where the picture does. At the default 25 fps, a two-frame drift is visible.
    2. Dialogue test. One speaker, one short line, a full-face framing. This is where most native-audio models fail first, and it fails obviously.
    3. Ambience test. An establishing shot with no foreground event. Judge whether the bed is plausible for the scene or generic room tone.
    4. Silence test. A prompt that should produce near-silence. Models that always emit a busy mix are hard to cut into an edited sequence.

    Run all four across several candidate models in one pass rather than one model at a time. In the agent, a single request can name a list of models and a fixed seed, so you get the same prompt rendered by each candidate instead of four separately-worded attempts you cannot fairly compare.

    The candidates worth putting in that eval

    The models that appear near the top of the audio-scoped board are the ones with a real single-pass audio path, and several are callable options in the catalog:

    • MiniMax H3 — unveiled 31 July 2026 at WAIC Shanghai, omni-modal, 5–15 second clips at 2K with native stereo audio in one pass. A separate reference-to-video variant in the same family takes multimodal references — images for subject and style, video clips for motion, audio clips cited in the prompt by order.
    • FLUX 3 — Black Forest Labs' first video model, announced 23 July 2026, one model jointly trained on image, video and audio, with synchronised sound up to 20 seconds.
    • Gemini Omni Flash — top of Arena's board and second on the audio-scoped board, which makes it the rare model that is not a board-shopping artifact.
    • Seedance 2.0 — ByteDance's line, present in the top four on both boards under slightly different entry names.

    Two names you will see quoted and should treat carefully. Wan 3.0 entered public beta on 6 August 2026 with 30-second clips to 1080p and audio in one pass, but its open-weights status is contested across sources and there is no confirmed checkpoint; treat it as closed until Alibaba says otherwise. Seedance 2.5 shipped on 31 July 2026 into Jimeng AI and Doubao Pro, China-first consumer surfaces, with no callable Ark endpoint at launch; developer API access followed roughly a week later. Shipped and buildable are separate milestones, and a board position never tells you which side of that gap a model is on.

    For a shortlist filtered by what actually runs rather than what ranks, models with audio is the catalog view.

    FAQ

    Is one of these boards more trustworthy than the other?

    No, and framing it that way is the mistake. They measure different things honestly. The audio-scoped board is the correct citation for sound-on work and the wrong one for silent B-roll. The failure mode is not a bad board, it is citing a board whose scope does not match your brief.

    Why does the same model appear under different names on the two boards?

    Because arenas list the specific build they tested. Arena carries dreamina-seedance-2.0-720p and veo-3.1-audio; a model family often has several entries at different resolutions or with audio on or off. Match on the exact entry name before you compare Elo across boards, or you will compare a 720p variant against a 1080p one and call the gap a quality difference.

    Should I use Elo to pick between two models that are 20 points apart?

    Not on its own. A 20-point gap means one wins the head-to-head slightly more often across a generic prompt mix. Read gaps, not values, and treat anything under roughly 50 points as a tie you should break with your own four-prompt eval on your own footage.

    Does a top rank mean I can actually use the model?

    No. As of mid-August 2026, Arena's top ten includes a model that has never shipped publicly and one whose API shuts on 24 September 2026. Check availability separately from rank — they are unrelated facts that leaderboards present side by side.