Reading the quality-per-credit report
A reading method for the quality-per-credit frontier that shortlists models for one specific job and still works after the next round of launches.
The default way to pick a model is to sort by price and take something near the top, on the theory that expensive means good. That theory is testable, and the quality-per-credit frontier report tests it: it plots third-party arena score against credits per job for every scored model in the catalog, and most of them turn out to be dominated — something else gives more score for the same credits, or the same score for fewer.
The report is not a ranking you read top to bottom. It is a chart you read sideways. Here is the method.
What is actually on the axes
Two things need pinning down before any row means anything.
The vertical axis is not Versely's opinion. The quality scores come from Artificial Analysis, an independent benchmarking site that runs public human-preference arenas for image, video and speech models. Versely syncs those arena ELO figures into the catalog and matches them to catalog models by name. Every row on the report shows the exact leaderboard entry a score was taken from, in a column labelled "Scored as". The scores are theirs; the credit prices are Versely's. If you have seen a different number for the same model somewhere else, check whether the other source was quoting a different arena — ELO ratings are only comparable within the board that produced them.
The horizontal axis is credits per job, and "per job" is load-bearing. It means the fewest credits one complete generation can cost, taken from the model's own price matrix. That is the only figure comparable across models that bill per second, per megapixel and flat per call — a headline credit figure is not, because "10 credits per second" and "40 credits for the whole clip" are not the same kind of number. A handful of scored speech models publish only a per-1,000-character rate with no complete-job minimum; those rows are labelled "headline rate" and are kept out of the charts and out of the frontier rather than being passed off as job prices.
The report covers five arenas: text-to-video, image-to-video, text-to-image, image editing and text-to-speech. Pick the one that matches your job before you read anything. "Best video model" is not a question the report answers, because text-to-video and image-to-video are separate boards with different pools.
Read the frontier, not the ranking
A model sits on the frontier when nothing in its arena scores higher at its price or below. Everything off the frontier is dominated by definition. Frontier rows are highlighted in every table, and they are the only rows worth reading first.
That collapses a long table into a short list very quickly. The rest of the ranking is useful for a different question — "is this specific model any good?" — but it is noise when the question is "what should I use for this job?"
The frontier list is ordered cheapest first, and reading it in that order is the whole technique:
- Start at the cheapest frontier row and note its score.
- Step to the next row up. Ask what the score gained, and what it cost in credits.
- Keep stepping until the gain stops being worth the step. That point is your knee, and it is personal — it depends on whether the output is a hero asset or one of forty b-roll cuts.
- Shortlist the two rows either side of the knee. Not one. You want a fallback that is a known quantity.
Almost everyone's knee is lower than they expect, because the frontier is steep at the bottom and flat at the top. The last few ELO points in an arena regularly cost several times what the first several hundred did.
Three checks before you trust a row
Check what the score was earned on. The "Scored as" column exists because the match is by name, and a leaderboard entry often covers a model family rather than the exact variant you are about to run. A Fast or Lite variant sitting under a family's score is not a lie, but it is not a measurement of that variant either. When the distinction matters, treat the score as a ceiling for the cheaper variants in the family.
Check whether the row declares a job cost. If the credit figure is labelled "headline rate", the model publishes no complete-job minimum and you are looking at a rate, not a price. Those rows are excluded from the frontier for exactly this reason.
Check what ELO cannot see. Arena score measures which output people preferred across broad prompt sets. It says nothing about your aspect ratio, your clip length, whether you need native audio, whether you can feed it a reference image, or how long it takes to come back. A model can be on the frontier and still be wrong for you because it does not sell the duration you need. That is what the individual pages in the model catalog and the use-case cuts under the /best rankings are for.
One related trap worth naming: a "#3" on one leaderboard and a "#3" on another are not the same claim. Rank positions belong to the board that produced them, and several models across video, image and audio hold a rank of 3 simultaneously on their own boards. Never compare rank numbers across arenas.
The cohort tables, and what they do not say
Under each arena, a second table groups models by release year and shows each cohort's best score today and its lowest job floor. It is tempting to read that as a history of the leaderboard. It is not, and the report says so before the first table.
The catalog stores each model's current arena score only. There is no score history. So the cohort view is constructed by grouping models by release date and reading today's scores across those groups. It answers "what is each generation of models worth now?" — not "what did the board look like in March?" Older models' scores keep moving as new arena votes come in, so a 2025 model's number on this page is a 2026 measurement of a 2025 release.
Read that way, the cohort table is still the most useful part of the report for planning, because it separates two things that get conflated. New releases mostly move the floor, not the ceiling: the interesting figure is usually not "the best 2026 model scores higher" but "the cheapest way to reach last year's best score is now this much less". Cohorts with fewer than three priced models are shown but not narrated, which is a deliberate refusal to draw a trend line through two dots.
A method that survives the next launch
The point of a reading method rather than a shortlist is that model launches invalidate shortlists roughly monthly. This sequence does not go stale:
- Name the job as an arena. Text-to-video, image-to-video, text-to-image, image editing or text-to-speech. If your job spans two, run the method twice.
- Read only the frontier rows, cheapest first, and find your knee.
- Take the two rows either side of it as your shortlist.
- Verify the constraints on the model pages — durations, aspect ratios, reference inputs, native audio. Drop anything that cannot do the job regardless of its score.
- Check the cohort table for the floor, not the ceiling. If the cheapest way to reach your required score dropped since you last looked, your knee moved.
- Re-run it when a launch lands. New models surface in what's new first, and a launch that lowers the floor changes step 2's answer even if it never touches the top of the board.
Steps 4 and 5 are where most of the value is, and they are the two people skip. The frontier tells you where quality stops being worth extra credits. It does not know your brief, and it never will. For the budget side of the same decision, the credits page explains what a job costs across the catalog, and best-value AI video model applies a similar logic to one content type with the use-case constraints already folded in.
FAQ
Does a higher arena ELO mean the model is better for my job?
Not automatically. Arena ELO measures which output people preferred across broad prompt sets. It says nothing about your specific style, aspect ratio, clip length or budget, and it cannot know whether the model even supports the thing you need. Use the report to find the price band where quality stops improving for you, then let the constraints decide between the two or three candidates that survive.
Why is a model I use missing from the report?
Two common reasons. Either it holds no current score in any of the five arenas the ranking sync covers, or it has no page in the published catalog — the report's pool is one row per model page, so every row is something you can open and run. A model can be perfectly good and simply not be on a public human-preference board.
Why are some credit figures a range and others a single number?
Because billing shape differs. A flat-priced model has one number; a model billed per second or by resolution band has a floor and a ceiling. The frontier is computed on the floor, since that is the only quantity comparable across shapes, so a wide range on a frontier row means the ceiling is a lot further away than the chart position suggests.
Should I just always pick the cheapest frontier model?
No, and that is the failure mode on the other side. The cheapest frontier row is the best score available at the bottom of the price range, which is not the same as being good enough for your brief. Walk up the frontier until the score gain stops mattering to you, then stop. Where that happens is a judgement about the work, not about the chart.