Single-model mastery vs routing a whole catalog
Knowing one model deeply beats a marginally better one you cannot steer. Two thresholds, job variety and output volume, for when that stops being true.
The case for using one model and ignoring the rest is stronger than the catalog-shopping crowd admits. A prompt vocabulary tuned against one model's specific failure modes is worth more than a marginally better model you cannot steer, and it is worth more than most people estimate because the value is invisible: it shows up as generations you did not have to run.
The case against is also real. A catalog of 296 models exists because no single model covers every mode, every duration, every aspect ratio and every price shape, and pretending otherwise means paying premium rates for jobs your favourite is bad at. Both arguments are correct. What is missing is the threshold, so here it is.
What mastery actually buys
Model familiarity is not a preference, it is accumulated calibration. Specifically:
- A prompt vocabulary that lands. You know which adjectives this model responds to and which it ignores, whether it wants camera direction before or after the subject, how much detail is productive before returns go negative. None of that transfers. It is genuinely per-model knowledge, which is why switching feels like a step backwards even when the new model benchmarks higher.
- Known failure modes. You know that this model loses hands at high motion, or drifts on the last second, or ignores the third clause. Knowing where a model breaks lets you write prompts that route around the break, which is a skill you cannot acquire from a spec sheet.
- A predictable retake rate. This is the one that shows up in the budget. If you land a usable clip in two attempts on a model you know and four on one you do not, the unfamiliar model has to be dramatically cheaper per job just to break even, and it usually is not.
- Reusable seeds and settings. A settings block that produced a good result is a reusable asset only as long as you stay on the model.
The corollary is that switching has a real cost, paid in bad generations, and that cost is not on any comparison page. Anyone telling you to switch because a new model ranks two places higher is quoting a benefit and omitting the invoice.
What routing buys
Mode coverage. This is the strongest argument and the least emotional one. Text-to-video, image-to-video, reference-to-video, video-to-video, extend, lipsync, motion control and first-last-frame are separate catalog entries, often inside the same model family, and no single entry does all of them. The moment your work includes a second mode, you are routing whether you call it that or not.
Price shape, which is not the same as price. Models bill on different units, and the units behave differently as the job grows. Among the models with published pages, roughly 33 price per second of output, 18 per megapixel, 10 per 1,000 characters and 25 flat per call, with a substantial group declaring no unit at all. Two models showing the same headline credit figure can land far apart on a 15-second render. Worse, on a good number of catalog records the headline figure does not even appear in the model's own option list, which is why the comparable quantity is the job cost, the fewest credits one complete generation can cost, and not the number on the card. The billing shapes index is a fifteen-minute read that pays for itself.
Ceilings. Duration ladders, resolution caps and aspect ratio sets differ, and a ceiling is binary. Veo 3.1 stops at 8 seconds. FLUX.3's video ladder runs to 20. If your shot needs 12, the quality argument never gets made.
Actual price spread. The stated job ranges are wide enough to matter across a batch. The catalog puts a Veo 3.1 job at 80 to 320 credits depending on settings, FLUX.3 text-to-video at 43 to 170, Sora 2 text-to-video at 40 to 200. On one hero shot that is noise. On four hundred variants it is the project.
The two thresholds
Threshold one: job variety. Count the distinct generation modes you actually used last month. Not models, modes.
- One or two modes: stay on one model per mode and go deep. Routing buys you almost nothing and costs you calibration.
- Three modes: borderline. Route only where your default has a hard ceiling.
- Four or more: route. At four modes you are already using several models; the only question is whether you are choosing them deliberately or by accident.
Threshold two: output volume. The arithmetic is a comparison, not a number, and it takes two minutes:
- Take the credit gap per job between your familiar default and the cheaper alternative, using job cost rather than headline figures.
- Multiply by monthly volume. That is what mastery is costing you.
- Estimate extra retakes per job on the unfamiliar model. Two is a realistic starting assumption for the first fortnight.
- Multiply extra retakes by the alternative's job cost, times volume. That is what routing is costing you.
Below roughly fifty finished assets a month, step 4 almost always exceeds step 2 and mastery wins. Above a few hundred, step 2 dominates and routing wins by a distance. The middle is genuinely a judgement call, and the tiebreaker is whether your work is repetitive: identical jobs amortise the learning cost of a new model quickly, varied jobs never do.
Note that the retake term decays. It is a one-off cost of adopting a model, not a recurring one, which is why the answer changes as a team matures rather than staying fixed.
The hybrid that actually works
Not one default. One default per mode.
Write down the modes you use, and next to each, one production model and one fallback. That is the entire system. It gives you the calibration benefits of mastery, because within any given mode you are still using one model repeatedly, and the coverage benefits of routing, because the mode determines the model rather than habit doing it.
Then add a challenger rule, because the failure mode of any default list is that it silently goes stale. A challenger replaces an incumbent only if it wins on your own prompt set, not on a leaderboard and not in a launch video. Run the same three prompts you always run, three seeds each, and score on prompt adherence, subject consistency and artifacts. If the challenger does not clearly win, the incumbent stays, and the entire evaluation cost you nine generations. There is a full protocol in A/B testing AI models on one prompt.
Re-run it quarterly, not per release. Watching every launch and re-evaluating each time is how you end up with no defaults at all.
Making routing cheap enough to actually do
The reason people do not route is friction, so remove it.
The agent takes a single prompt and dispatches it across several named models in one request, which is the right shape for a bake-off because the model becomes the only variable. Same prompt, same seed where supported, same aspect ratio, different models, side by side. That is a five-minute exercise rather than an afternoon of tab-switching. Ask it to estimate the cost before dispatch and it will return per-item and total credits against your balance using the same pricing logic as the real charge.
For the price-versus-quality question specifically, the quality-per-credit report does the work already. It plots the models that nothing else beats on both price and score, using Artificial Analysis ratings synced into the catalog alongside each model's own job cost in credits. One caveat worth carrying: a model's rank belongs to the specific leaderboard it was earned on, and a model can sit high on image-to-video and much lower on text-to-video. Read the arena name, not just the number. If you want a shorter answer, best-value AI video model is the shortlist version, and the head-to-head pages cover specific pairs.
FAQ
Does mastery still matter when models change every few months?
Yes, more than it did. Version churn means the calibration you build is on a moving target, which sounds like an argument against depth until you notice that point releases mostly preserve a model's character. A prompt vocabulary tuned to a family survives a version bump far better than it survives a jump to a different family. Depth within a family is durable; depth in a specific build is not.
How many models should a small team actually run?
One per mode you use regularly, plus a fallback each, which for most teams means four to eight entries total. More than that and nobody is calibrated on anything. Fewer and you are paying premium rates for jobs outside your default's strengths.
Is the cheapest model per job always the right routing target?
No, because job cost does not include retakes, and a model that needs three attempts at a low rate is more expensive than one that needs one at a high rate. Route on expected credits per usable output, which means you have to measure your own hit rate per model. Almost nobody does this and it is the single highest-value number a team can track. How credits work covers the billing side of that calculation.
What is the fastest way to tell if my default is wrong for a job?
Run the job on your default and on one alternative from the same mode, same prompt, three seeds each. If your default wins, you have spent six generations confirming it and you can stop asking. If it loses clearly, you have found a mode where your habit is costing you. Either result is worth the six generations, which is why this beats reading another comparison, including this one.