Evaluating a New AI Model: The 30-Minute Checklist
A 30-minute checklist for evaluating new AI video models: fixed prompt sets, stress tests, cost-per-keeper math, and a clear adopt-or-skip call.
A new model drops, the launch reel looks incredible, and your feed fills with hand-picked masterpieces. None of that tells you the only thing that matters: whether this model beats your current one on your content. Launch demos are the best 0.1% of outputs, chosen by the people selling the model. Your evaluation has to be the opposite — fixed prompts, honest counting, and a decision at the end.
The good news: a rigorous-enough evaluation takes about 30 minutes and a modest pile of credits. Here's the checklist, in the order that saves the most time.
Minute 0–5: Disqualify before you generate
Half of new models can be ruled out for your use without spending a credit. Check the spec sheet against your actual needs:
- Aspect ratios. You're a shorts creator and it's 16:9 only? Done — skip it.
- Duration limits. Your format needs 10-second shots and it caps at 5 with no extend? Skip, or note it for a different job.
- Input modes. If your pipeline is image-to-video and the release is text-to-video only, evaluate it later when the i2v variant ships (it usually does).
- Audio. If dialogue scenes are your format, native audio support is a binary gate.
- Cost tier. Note credits per generation now — you'll need it for the math in the last step.
Model pages on /models carry these specs in comparable form, which turns this step into a two-minute scan. Only models that pass the gates earn generation time.
Minute 5–20: Run your fixed prompt set
The core of the evaluation is a standard prompt set you reuse for every model — this is the single habit that separates signal from vibes. If you don't have one yet, build it once from your last month of real work: pull the 5–6 prompts that represent what you actually make, not what looks cool.
A balanced set covers:
- Your bread-and-butter shot (e.g., a product on a surface, slow push-in) — the one you'd run 50 times a month.
- A human shot — faces and hands are where models diverge most.
- A motion stress test — running, liquids, cloth, particles; physics failures live here.
- A text-in-scene shot — signage or packaging, a classic weak point.
- Your style prompt — whatever aesthetic your brand leans on.
- One weird one — an unusual combination that punishes models that only memorized common scenes.
Run each prompt 2–3 times (single samples lie — variance between takes is itself a metric you're measuring), with settings matched to your current model's output: same aspect, same duration where possible. Where the model exposes a seed, note it so promising results can be revisited; the methodology in A/B testing AI models with one prompt covers keeping comparisons honest.
While they render, resist the launch-thread temptation. Other people's cherry-picks will anchor you before your own data arrives.
Minute 20–25: Score like an editor, not a fan
Review outputs side by side with your incumbent model's results on the same prompts (this is why the fixed set pays off — you already have those). Score each prompt 0–2: 0 = unusable, 1 = usable with fixes, 2 = publishable as-is. Then compute the only number that matters:
Keeper rate — publishable outputs ÷ total generations.
Keeper rate matters more than peak quality because it is your real cost. A model with stunning best-case output but a 1-in-6 keeper rate costs six generations per usable clip; a slightly-less-stunning model that keeps 1-in-2 is three times cheaper in practice and infinitely cheaper in iteration time.
| Metric | How to measure | Why it matters |
|---|---|---|
| Keeper rate | Publishable ÷ total generations | Your true cost multiplier |
| Consistency | Score spread across takes of one prompt | Predictability = plannable workflows |
| Prompt adherence | Did it render what you asked, or its own idea? | Determines revision loops |
| Motion quality | Physics, hands, temporal stability | Where video models actually differ |
| Cost per keeper | (Credits per gen) ÷ keeper rate | The comparison number across models |
| Render speed | Wall-clock per generation | Compounds brutally at volume |
Cost per keeper is the line that decides things: credits per generation divided by keeper rate. A model that's 30% pricier per generation but keeps twice as many outputs is dramatically cheaper. Run this arithmetic before any adoption decision — it reverses first impressions surprisingly often.
Minute 25–30: Make one of three calls
Force a decision — an evaluation without a decision is just entertainment:
- Adopt — it beats your incumbent on cost-per-keeper for a job you do weekly. Swap it into that job (not everything at once), monitor for a week, then expand.
- Slot — it doesn't dethrone your main model but wins a niche: the fast drafts, the text-heavy shots, the human close-ups. Multi-model workflows are the norm now, and most "losing" models still win somewhere.
- Skip — note why in one line ("weak hands, 0.2 keeper on people shots, revisit at next version") so the next release from this family starts with context.
Two cautions at the decision point. First, don't overweight day-one performance in either direction — serving infrastructure often improves in the weeks after launch, and occasionally quality is quietly tuned down under load, so re-run the set once before betting a big project on a new adoptee. Second, cross-check your conclusion against community consensus after forming your own: live ELO standings (explained in reading the model leaderboard) tell you how the model performs across thousands of matchups, which catches blind spots in a six-prompt set — but your set weighs your niche in a way no global ranking can.
Keep the flywheel: your eval set compounds
The first evaluation takes 30 minutes plus setup. Every subsequent one takes 30 minutes flat, and the results become comparable across time — a private benchmark of every model you've ever tested on the exact work you do. Store the outputs in labeled folders, keep the scores in one spreadsheet, and retire prompts from the set only when your content itself changes. Within a few releases you'll know your niche's model landscape better than any reviewer covering models in general — because nobody else is benchmarking your content.
FAQ
How many credits should I budget for evaluating a new model?
Around 12–18 generations covers a six-prompt set at two to three takes each — modest enough that free daily credits plus a small balance handles it. Skipping repetition to save credits is a false economy: single samples can't reveal keeper rate, which is the entire point of the exercise.
Why not just trust launch demos and reviews?
Launch demos are selected best cases, and reviewers test generic prompts that may not overlap your niche at all. A model can be genuinely great on average and still lose to your incumbent on your specific formats — only a fixed personal prompt set catches that.
What's a good keeper rate for an AI video model?
It varies hugely by prompt difficulty, which is why you compare against your incumbent on identical prompts rather than an absolute bar. As rough intuition: bread-and-butter shots should keep at one-in-two or better on a model worth adopting, while stress-test prompts keeping one-in-three is respectable anywhere.
Should I test a new model with its recommended settings or my usual ones?
Start with your usual settings so the comparison against your incumbent is fair, then give promising models a second pass with their recommended settings to see their ceiling. Adopting a model that only performs under settings your workflow can't use is a common trap.
How often should I re-evaluate models I already skipped?
At the next major version, or when the family ships a variant targeting your disqualifying reason — a fast tier if cost killed it, an i2v mode if inputs did. Your one-line skip notes make these re-checks nearly free, and model families often fix their weaknesses within a generation or two.
Build the habit where the models already live: Versely's AI video generator puts 60+ models behind one prompt box, so running your eval set against this week's newcomer is a dropdown change — and the live rankings on /models tell you which newcomer deserves the 30 minutes.