Filtering a huge model catalog down to eight
An elimination protocol that takes 296 models to 8 in five ordered gates: input modality, duration ceiling, audio, licence, credit band. Refill quarterly.
A shortlist built by addition never gets short. You read a launch post, add a model, read another, add another, and three weeks later you have nineteen candidates and no way to choose. The catalog carries 296 models. Adding your way to a working set from that number does not converge.
Elimination does. Five gates, applied in a fixed order, take the whole catalog to a testable handful in about ten minutes, and the order is what does the work: each gate is cheap to evaluate and each one is placed so it removes the largest possible block of models before the next, more expensive judgement has to be made. Here is the protocol with a worked example carried all the way through.
Why the order is the method
Every gate has a cost to apply. Checking a model's input modality is a field lookup. Checking whether its licence terms cover your client's territory is a document read. Check licence first and you will read documents for models that were never going to survive the duration gate.
So the gates are ordered by cost, cheapest first, and by cut size, largest first. Modality is both the cheapest to check and the most brutal, which is why it goes first. Credit band is last because it is the only gate you might reasonably relax if the survivors are too few, and relaxing the last gate does not force you to redo any of the others.
The one rule that matters: never skip ahead to the gate you find interesting. Everyone wants to start with quality, which is the most expensive judgement available and the one that requires generations to settle. Quality is not a gate. It is what you spend the shortlist testing.
The five gates, in order
Gate 1 — input modality. What do you have in hand, and what has to come out? Not "video model" but the exact mode: text-to-video, image-to-video, reference-to-video, video-to-video, extend. These are separate catalog entries even inside one model family, and picking the wrong entry is the single most common wasted dispatch. Browse by category rather than by brand.
Gate 2 — duration ceiling. Your longest required single take, not your finished runtime. A 45-second ad assembled from six shots needs a model whose ceiling clears your longest shot, not 45 seconds. Take the ceiling from the model's duration ladder and check the granularity while you are there, because a model that offers 5s, 10s and 15s and nothing between will make you buy seconds you trim away.
Gate 3 — audio. Binary and unforgiving. If the deliverable needs speech or synchronised sound out of the generator, models without an audio pass are gone regardless of how good they look. If you are laying sound in the edit anyway, skip this gate rather than pretending it removes anyone.
Gate 4 — licence terms. The gate the catalog cannot run for you. More on this below.
Gate 5 — credit band. A ceiling on the published credit figure for the model. Applied last, and applied loosely, because for a large share of video models the published figure is a per-second or per-resolution rate rather than a flat price, so the band is a sorting proxy rather than a quote. The billing shapes index is worth reading once: per_second, resolution_based_per_second and base_plus_per_second behave very differently as your clip gets longer, and two models with the same headline figure can land far apart on a 15-second render.
The worked example
The brief: a product ad built from a supplied hero photo, longest single take 12 seconds, dialogue and sound needed out of the generator, commercial use, and a cost profile that lets the team iterate rather than budget each dispatch.
| Gate | Applied | Models remaining |
|---|---|---|
| Start | Whole catalog | 296 |
| Content type | Video | 146 |
| 1. Input modality | Image-to-video | 52 |
| 2. Duration ceiling | 12s or longer | 27 |
| 3. Audio | Native audio pass | 17 |
| 4. Licence | Manual read of survivors | 17 |
| 5. Credit band | 20 credits or under on the card | 8 |
The eight that come out: Sora 2 image-to-video, Seedance 1.5 Pro image-to-video, Seedance 2.0 Mini image-to-video, Grok Imagine image-to-video, Wan 2.6 image-to-video, Vidu Q3 image-to-video, FLUX.3 first-last-frame and FLUX.3 image-to-video.
Two things are worth noticing. The duration gate did the heaviest lifting, cutting 52 to 27, because most image-to-video models stop at 10 seconds. The audio gate cut nearly as hard, 27 to 17, because native sound is still a minority capability. If your brief had not needed audio you would be choosing from a very different seventeen, which is the argument for writing the brief before opening the catalog rather than after.
The band at Gate 5 is a dial, not a verdict. Loosen it to 30 credits and you get ten survivors including Seedance 2.0's mid tier and Sora 2's storyboard mode. Loosen it further and the Kling V3 and O3 tiers arrive. Tighten it and you drop to five. Adjust the last gate until the survivor count sits between roughly six and ten, which is the range a small team can actually test in a session.
The gate the catalog cannot run for you
Licence sits at position four for a reason: it is the only gate that requires reading a document, and by that point you have 17 documents to read instead of 296. What you are checking is not "is this allowed" in the abstract but four specific things against your actual deliverable:
- Commercial use for your output type. Ad creative, broadcast and paid placement are sometimes carved out separately from general commercial use.
- Territory. Some model licences exclude specific regions, and a client with distribution in an excluded market is a problem you want to find before the shoot, not after.
- Attribution and disclosure. Whether output must be labelled, and whether that label is compatible with the placement.
- Redistribution and resale. Relevant the moment you are handing masters to a client who will license them onward.
None of that is a catalog field, and none of it is stable enough to hardcode into a shortlist you reuse for a year. Read the terms for the survivors, note the answers next to the model name, and re-check when a version number moves. The usage rights entry covers the vocabulary if the terms language is unfamiliar.
What the eight are for, and when to refill
The shortlist is not a ranking. Nothing about the funnel says Vidu Q3 is better than Wan 2.6, because none of the five gates measured quality. What the funnel produced is a set of models that are all eligible, which is the precondition for a comparison that means anything.
From there, run one fixed prompt through all eight and judge output. The agent takes a single prompt and fans it across several named models in one request, which is the right shape for this because the whole point is that the model is the only variable. Score on prompt adherence, subject identity and motion artifacts, keep the top two or three, and put the rest in a note with the date. Eight is a testing set; two is a production default plus a fallback.
Then leave it alone. The failure mode after a first pass is watching every release and re-running the protocol each time something ships, which is how you end up back at nineteen candidates. Set a quarterly refill instead: re-run the same five gates against the current catalog, note which models are new to the survivor set, and test only the new entrants against your existing default. A challenger has to beat the incumbent on your own prompt set to replace it, not just look better in a launch video.
The gates themselves should barely change between quarters. If they are changing, the brief changed, and that is a different exercise: picking the generation mode before the model is the piece to re-read when the job itself has moved.
FAQ
Why not just start from a "best model" list?
Because a best-of list is answering someone else's brief. A ranked list optimises for a general notion of output quality, which is genuinely useful once your constraints are satisfied and completely useless before that. A model can be the highest-ranked video generator available and still be disqualified by a 10-second ceiling when your shot needs 12. Constraints first, quality second.
What if a gate leaves me with zero models?
Then a constraint in the brief is not achievable and you need to know that now rather than three days in. Work backwards through the gates and identify which one produced the empty set. If it was duration, the answer is usually to split the shot rather than find a longer model. If it was audio, the answer is a separate sound pass in the edit. If it was the credit band, the answer is a decision about budget, and at least it is now a decision with a number attached.
Does this work for image models?
The same order works with two substitutions: duration ceiling becomes output resolution, and audio becomes whichever binary capability the job depends on, usually text rendering or reference support. Modality still goes first, because text-to-image, image-to-image and instruction editing are separate entries and picking the wrong one wastes the dispatch exactly the same way.
How long should the whole protocol take?
Ten minutes for gates one through three and five, since those are all field lookups you can run from the model catalog. Gate four takes real time, maybe an hour across seventeen models, and it is the one people skip. It is also the only gate that can produce a legal problem rather than a creative one, so it is the wrong place to save forty minutes.
Run the funnel once properly and the output is a shortlist you can defend to a client, a producer or a finance team, because every model that is not on it was removed by a stated rule rather than a preference.