Comparisons

    A scoring rubric for comparing video models

    A weighted five-axis rubric with written 1-5 anchors for every point, so a model bake-off produces a number you can defend a week later.

    Versely Team9 min read

    Most model bake-offs don't survive a second opinion, because "this one looks better" is not a measurement. It's a preference expressed once, by one person, in one mood, and it evaporates the moment someone asks why. The fix isn't more taste. It's writing down what each score means before you look at any output, so a 3 on Tuesday is the same 3 on Friday and the same 3 for the person next to you.

    Here's a rubric that does that: five axes, weighted, with a written anchor at every point on the scale.

    Why weighted, and why anchored

    Weighted, because the axes are not equally load-bearing. A model producing gorgeous, artefact-free motion while ignoring half your brief is not 80% as good as one that does what you asked — it's unusable, and an unweighted average will rank it near the top anyway. Weighting encodes that a brief-following failure costs more than a slightly soft hand.

    Anchored, because a bare 1-5 scale drifts. Unanchored scoring is a vibe with a number stapled to it: raters use different parts of the range, and the same rater's "4" migrates over a session, so by clip twelve your early and late scores are on different scales. A written sentence describing what a 2 looks like is the mechanism that makes the number portable.

    One convention that saves confusion later: score every axis so 5 is always good. Two of the five below are defect axes, where more of the thing is worse. Reverse them at the anchor level, not with a minus sign in the arithmetic, or someone will sum them wrong within a month.

    The five axes

    1. Prompt adherence — weight 30. Did it make the thing you asked for? See prompt adherence; this axis is instruction-following only, not beauty.

    • 1 — Ignores the brief. Wrong subject, wrong setting, or an entirely unrequested composition.
    • 2 — Subject roughly right; two or more explicit instructions dropped (named action, camera move, time of day).
    • 3 — Subject and setting right, one explicit instruction dropped or substituted.
    • 4 — Everything explicit is present; a nuance of tone, framing or pacing is off.
    • 5 — Every explicit instruction present, implied constraints respected without being stated.

    2. Subject consistency — weight 25. Does the subject stay the same subject throughout? This is temporal consistency applied to identity, and it decides whether a clip can sit next to another in a cut.

    • 1 — Identity breaks on screen. Face, garment or product morphs mid-shot.
    • 2 — Recognisable drift: hair, clothing detail or logo changes materially first second to last.
    • 3 — Holds through simple motion, breaks on a turn, occlusion or scale change.
    • 4 — Holds throughout; one small detail (a strap, a button, a background object) shifts.
    • 5 — Frame-one subject is frame-last subject at every level of detail.

    3. Motion artefacts — weight 20. Reverse-scored: 5 means fewest defects. Limbs, physics, warping, flicker — the tells that make a clip read as generated.

    • 1 — Immediately disqualifying. Extra or fused limbs, objects through solids, gross warping.
    • 2 — Obvious at full speed. A hand that reforms, a gait that never touches the ground.
    • 3 — Visible on a second viewing or in one region. Survives a scroll, fails a pause.
    • 4 — Visible only frame-stepping. A single soft frame at a transition point.
    • 5 — Nothing found at frame-step across the whole clip.

    4. Audio sync — weight 15. Only for models generating native audio. Covers lipsync where there's dialogue, event sync where there isn't — the footstep landing when the foot lands.

    • 1 — Audio unrelated to picture, or mouth movement with no correspondence to sound.
    • 2 — Correspondence exists but drifts, ending visibly out.
    • 3 — Broadly in sync; phonemes or impacts late or early enough to notice.
    • 4 — In sync throughout; timbre or ambience doesn't match the scene.
    • 5 — Sync holds at frame level and the sound belongs to the space.

    5. Seed variance — weight 10. Run the same prompt at three seeds and score how far apart the results land. Reverse-scored for production work: 5 means predictable.

    • 1 — Three seeds, three unrelated videos. No stable reading of the prompt.
    • 2 — Same subject, wildly different composition, lighting and pacing each time.
    • 3 — Same subject and composition; quality swings enough that one of three is unusable.
    • 4 — Consistent output; small variation in framing or performance.
    • 5 — Three near-siblings. What you got once is what you'll get again.

    Score each axis 1-5, multiply by weight, sum, divide by 100. One number between 1 and 5.

    The scorecard

    Axis Weight Model A Model B Model C
    Prompt adherence 30 5 3 4
    Subject consistency 25 3 5 4
    Motion artefacts 20 3 4 4
    Audio sync 15 2 4 3
    Seed variance 10 4 2 4
    Weighted score 100 3.55 3.75 3.85

    Those are illustrative numbers, not measurements — the point is the shape of the result. Model A wins outright on the axis everyone talks about and finishes last. Model C wins nothing and ranks first, because it has no weak axis. That pattern is the argument for the rubric existing: a weighted composite rewards absence of failure, which is what production needs, over a single spectacular strength, which is what a demo reel is made of.

    Handling missing axes properly. If a model has no audio, do not score audio sync 1. Treating an absent capability as a defect is a category error that quietly ranks every silent model last. Drop the axis and renormalise the remaining weights to 100 — prompt adherence 35, subject consistency 29, motion artefacts 24, seed variance 12 — and note on the scorecard that you did, so the two totals aren't compared as though they came from the same instrument.

    Running it without contaminating it

    Four controls, in order of how much damage skipping them does:

    1. Score blind. Strip model names first. Reputation moves scores more than any other single factor — a named frontier model gets half a point for free from almost everybody.
    2. Fix every variable except the model. Same prompt text, same duration, same aspect ratio, same resolution. Versely's default frame rate is 25 fps, so that holds unless you override it. One prompt across several named models in a single request is what generating a video from text is for — name the candidates and they go out together, with no chance of a retyped prompt drifting. The one thing you can't hold fixed is the seed: seeds don't transfer between models, so seed 42 on one has nothing in common with seed 42 on another. Lock it within a model, for axis five; across models it buys nothing.
    3. Score in axis order, not clip order. Score every clip's prompt adherence, then every clip's subject consistency. Scoring one clip fully before moving on lets your impression of axis one leak into axis three. Axis-first is the cheapest halo-effect control there is.
    4. Frame-step the artefact axis. Anchors 4 and 5 are not distinguishable at full speed by design. If you aren't stepping frames you're scoring a 3-point scale and calling it a 5-point one. Analysing a video is a reasonable first pass on a longer clip.

    To review a side-by-side strip rather than switching tabs, the editor will stitch the candidates into one timeline, and its 480p preview pass is free — with a short per-user cooldown between previews — so building the comparison reel doesn't eat the budget meant for the candidates. The generations still cost credits; the preview render of the assembly doesn't.

    Where public leaderboards fit

    Use them as a prior, not a result. They measure pairwise human preference across a general prompt distribution — a different question from "does this follow instructions on the handful of shot types my client keeps asking for."

    Be specific about which board you cite. Arena (arena.ai), which was LMArena until January 2026, and Artificial Analysis (artificialanalysis.ai) are unrelated organisations running separate vote pools on separate scales, so a number from one is not comparable to a number from the other. Their mid-2026 text-to-video orderings mostly agree, but the exceptions are wide: a model can top one board and not appear in the other's top ten, often because the two added it at different times rather than because either is wrong. Versely's catalog ELO scores are synced from Artificial Analysis and matched by model name, which is what the quality-per-credit report plots credit price against. Good for shortlisting, bad as a substitute for scoring your own footage. When a vendor claims "#1 model," the first question is which board.

    FAQ

    How many clips per model do I need?

    Three seeds on each of several prompts, minimum. One clip per model measures luck. Seed variance is unscoreable below three, and it's the axis that most often explains why a model that demoed brilliantly disappoints in production.

    Should I change the weights?

    Yes if your work justifies it, and once — before scoring, written down. Silent product shots can push subject consistency to 35 and drop audio entirely. What breaks the method is adjusting weights after seeing scores, which is how you construct a rubric that ranks the model you already liked first.

    Can one person run this credibly?

    Yes, provided the anchors are written before the first clip. Anchors are what make a single rater reproducible; without them a single rater is a preference with extra steps.

    Does a higher composite mean I should switch models?

    Not on its own. It says which model produces better output per attempt, not which is a better deal per finished shot — that depends on credit price and reroll frequency, a separate calculation that runs on top of this one.

    Run it once a quarter and on every new release. It fits on an index card, which is the only reason it survives a production week.