Quality per credit beats arena rank for repeat jobs
A weekly series should pick the model that wins the job at the credit floor, not the Elo leader of the week.
Arena Elo is a blind vote on an eight-second clip. It is a real measurement. It is not a buying guide for a show you have to ship every Thursday.
A weekly series should pick the model that wins the job at the credit floor, then leave it there. Chasing the Artificial Analysis leader of the week is how a series changes grain, motion cadence and skin every episode while the credit burn goes up. Arena Elo is not the buying guide is the job-table version. This post is the repeat-job version: quality per credit, on purpose.
The cost function is not Elo
Elo answers: in a head-to-head, which eight seconds look better to a rater who has never seen your brand. A series answers: can this model lock the character, hold the product, speak if the shot talks, and do that fifty times without the look drifting and without the bill becoming the story.
Those are different tests. Omni Flash can lead a text-to-video-with-audio board and still be the wrong pick for a talking-head ad that needs Veo's always-on dialogue. Seedance can own image-to-video and still be waste on a landscape plate you will grade later. Kling can win a motion week and still be the wrong unit for a pack shot that has to match a still.
Repeat work has a fourth variable the board ignores: you already picked. Episode two that matches episode one is worth more than episode two that is "better" in a vacuum. Switching to the new Elo leader is a look change. Treat it like a rebrand, not like a default.
Read the frontier, not the podium
Quality per credit plots arena score against Versely credits per job, on the published catalog, with scores from Artificial Analysis and prices from our matrix. The useful object is the Pareto frontier: models nobody else beats on both score and credit cost. Everything above and to the right of a frontier row is paying extra for a win you can get cheaper, or paying the same for a worse score.
Use it like this:
- Name the job. Text-to-video, image-to-video, stills — the report is split by arena on purpose.
- Find the cheapest row that still clears your bar, not the top Elo. Your bar is "identity holds," "type is legible," "audio is usable," not "won Tuesday."
- Pin that model for the series. Write it next to the palette.
- Revisit when the job changes, or when a frontier row undercuts your pin on the same job at a cost you care about. Do not revisit because a blog roundup crowned a new king.
Credits only. The report does not publish a dollar-per-clip. What it costs is the formula book when you need to know whether this job is per-second, per-character, or flat per call. A weekly series that mixes talking shots, silent B-roll and TTS is three formulas. Rank each one. A single "best model" is how you overpay two of the three.
What "wins the job" means in practice
Identity lock. If the still is the product, I2V or R2V on a model that holds references beats a higher-Elo text-to-video that invents a new face. Retries are the real cost. A cheaper model that hits once is cheaper than a podium model you generate four times.
Audio. If the mouth is in frame, native audio on a dialogue specialist beats a silent Elo winner plus a TTS slap. If you need a cloned founder voice across thirty episodes, TTS wins and native audio is a fight with the mouth. Do not mix both on one mouth.
Length. A thirty-second identity pass and eight-second stitches are different bills and different continuity taxes. Pick the object that matches the cut, then pick the model that can emit that object without a heroics pipeline.
Control. Runway-class camera tools and first-last-frame turns are not arena jobs. If the brief is a specific move, Elo is almost off-topic.
| Repeat job | Rank this | Ignore this |
|---|---|---|
| Weekly explainer, same host still | I2V lock at the credit floor | T2V Elo of the week |
| Talking product spot | Always-on dialogue, then price | Silent beauty clips |
| Hook pack volume | Distinct openings, pack billing | Cinematic 16:9 arena clips |
| Graded landscape plates | HDR / bit depth, then price | Native audio |
The best pages already rank by job. Use them when you are choosing a pin. Use the frontier report when you are arguing with someone who walked in with a leaderboard screenshot.
Pin, then stop
Write the model id in the series bible so routing does not swap families. When a new model lands on the frontier, bake it off on one episode — same stills, same prompt, same duration — and keep it only if it wins your checks at a cost you want.
Quality per credit is not a dare to use the cheapest row. The cheapest row that fails identity is the most expensive row you will run this month. The frontier is how you see that before you feel it.
FAQ
Should I switch models every time the arena updates?
No. Arena snapshots move. A series look should not. Switch when the job changed, the pin failed a locked check, or a frontier row beats your pin on the same job at a cost that matters. Weekly Elo is a newsletter, not a production schedule.
Where do the Elo numbers on the report come from?
Artificial Analysis public arenas, matched to catalog models by name. Prices are Versely credits for a comparable job, not provider USD. The report says so on the page. Do not paste a screenshot of someone else's board into a budget thread as if it were our matrix.
What if the Elo leader is also the cheapest that wins my job?
Then you are already on the frontier and you can stop reading. Pin it. The failure case this post is for is the other one: paying podium prices for a job a cheaper row already clears.
How do I compare a per-second model to a flat-per-job model?
You do not, not on one number. Cost exists because the arithmetic differs. Compare them inside one job shape — eight seconds of talking video against eight seconds of talking video — then look at quality per credit inside that shape.