Regression-test your style when models change
A newer model quietly shifts colour, motion and faces across a look you already shipped. The golden-set pass that catches it before a client does.
Nobody evaluates a replacement model on the work they already shipped. They evaluate it on new prompts, where there is nothing to compare against, and every output looks like an improvement because there is no prior sitting next to it. Two weeks later a client asks why the founder's skin looks different in the new batch.
That is the whole failure mode. A model swap is not a quality question, it is a continuity question, and the two are measured differently. A model can beat your incumbent on every axis a public board scores and still break a look you spent four months establishing, because "better" and "the same" are unrelated properties.
The three things that move without telling you
Across a swap, most of what changes is invisible in isolation and obvious in a sequence.
Colour rendition. Contrast curve, saturation defaults, how highlights roll off, and — the one that gets noticed — skin tone in shadow. A model that renders a shade cooler is fine on its own and wrong the moment it cuts against six months of previous footage.
Motion. How literally a camera instruction is taken, how much motion blur appears, whether a "slow push in" is slow. Cadence matters here too: the default frame rate for video generation is 25 fps, and a timeline mixing sources that disagree about cadence produces judder you will blame on the wrong thing.
Faces. Identity drift across a set, eye-line, teeth, hairline, and apparent age. This is the axis with the least prompt-side recourse, which is why it belongs in the test rather than in the post-mortem.
Underneath all three sits a fourth: prompt interpretation. The same words weight differently. A term that reliably produced a look on the incumbent may be near-inert on the replacement, so a regression can appear as a style failure when it is really a vocabulary failure.
Build a golden set from shipped work
A golden set is not a capability test. A 20-prompt suite for testing any video model exists to find a stranger's weaknesses in the abstract; a golden set exists to prove that your look survived. It is drawn entirely from prompts you have already run in production, which means it is different for every shop and useless to anyone else.
Sixteen slots is enough. Fewer than twelve and one bad slot swings the read; more than twenty and nobody re-runs it.
| Slot | What it pins down | Count |
|---|---|---|
| Hero face, mid shot | Identity and skin tone in your standard light | 3 |
| Hero face, close | Eyes, teeth, hairline, micro-expression | 2 |
| Product on surface | Colour accuracy, material, label legibility | 3 |
| Signature camera move | How literally motion instructions are read | 2 |
| Signature lighting look | Contrast and highlight rolloff | 2 |
| Caption-safe frame | Headroom and composition where text lands | 2 |
| Known-hard shot | Your worst historic failure, kept deliberately | 2 |
Three rules make the set worth keeping. Freeze the prompt text verbatim, including the parts you suspect do nothing. Freeze the reference images, since changing a reference changes the experiment. And do not expect a seed to carry across — seeds are model-local, so a matching number on a different model means nothing.
Run the pass
- Re-run the incumbent first, the same day. This is the step people skip and it is the one that makes the result valid. Providers ship silent updates; last quarter's archived renders are not a control, they are a memory.
- Run the candidate on identical inputs. Same prompts, same references, same aspect ratio, same duration.
- Strip filenames and shuffle. Blind scoring, in pairs, in a random order. You know which model you want to win and it will show up in your scores otherwise.
- Score three axes only. Colour, motion, face. Prompt adherence, speed and cost are real questions but they are not this question, and folding them in turns a continuity check into a general bake-off. That job belongs to a weighted scoring rubric.
- Log first-pass usable, not best-of-n. Best-of-n hides exactly the degradation you are testing for. Usable rate is the benchmark nobody publishes covers why that number reorders rankings.
- Report per-slot deltas, never an average. An average across sixteen slots will happily conceal that every close-up broke.
Running the incumbent and the candidate on one prompt is less tedious than it sounds, because the agent chat can fan a single prompt across several named models in one request — the same input, dispatched to each, results side by side. Cost is visible in credits before you confirm, which keeps a sixteen-slot double pass from being a surprise.
Reading the result
The useful output of the pass is not a winner. It is a map of where the two models disagree, and each region of that map implies a different response.
| Finding | What it usually means | Response |
|---|---|---|
| One slot regressed, rest flat | Model-specific weakness, not a style break | Route that shot type to the incumbent |
| Colour shifted across every slot | Global grade difference | Often correctable downstream in the edit |
| Motion literal-mindedness changed | Prompt vocabulary weights differently | Re-tune phrasing, then re-run the pass |
| Faces drifted within a set | Consistency mechanism is not transferring | Change the mechanism, not the prompt |
| Everything improved, nothing matched | A genuinely different look | Treat as a rebrand, not an upgrade |
The face row is the one that costs money if you get it wrong. Prompt edits rarely recover identity drift; the fix is at the mechanism level, and four consistency mechanisms and when each fails is the map of which lever to reach for. If your look is currently held by a style reference that the new model reads differently, style references versus fine-tunes is the decision behind it. And if drift shows up within a single set rather than between models, the problem predates the swap and the swap is not the thing to fix.
"Everything improved, nothing matched" deserves its own note, because it is the outcome teams handle worst. A better look that is not your look is still a discontinuity to your audience. Either adopt it deliberately, with a cutover date and a refreshed set of brand references, or do not adopt it. What fails is drifting into it one campaign at a time.
When to run it
Four triggers, and only four, or this becomes a standing tax nobody pays.
- A provider sunsets something you depend on. Forced migrations are the highest-risk case because the deadline is not yours. The Sora API sunset migration plan is what one of those looks like with a date attached.
- A new version of a model you already use ships. Point releases are the sneakiest, because the name barely changes and nobody thinks to re-test.
- A price or capacity change tempts a swap. Run the pass before the spreadsheet decides.
- Quarterly, for the two or three models carrying your main look. Silent updates are real; a scheduled control run is how you find them.
Keep the outputs. A golden set with eighteen months of dated results behind it stops being a test and becomes a record of what your look actually was on any given date, which is the thing you will want when a client insists something changed and you need to know whether they are right.
FAQ
Can I skip re-running the incumbent to save credits?
You can, and the result stops being evidence. Comparing today's candidate to renders made months ago conflates a model change with every update the incumbent has had since. If budget is the constraint, cut the set to twelve slots and keep the control — a smaller honest test beats a larger broken one.
The new model is better on average but worse on faces. Now what?
Split the routing rather than picking a side. Faces stay on the incumbent, everything else moves, and the timeline reconciles the two. That is a normal steady state, not a failure to decide, and choosing the model yourself versus letting the agent route covers where the decision should sit.
Does this replace a full model comparison?
No. This asks one narrow question — did my established look survive — and is deliberately blind to speed, cost and range. Run it alongside a broader evaluation, or after one has produced a shortlist of two or three candidates worth the double pass.
How do I test the edit side rather than the generation side?
Rebuild one shipped cut with the new clips on the existing timeline. Because the editor is EDL-based, the timeline stays live and the swap is a re-render rather than a rebuild. Iterating there runs through preview: true, which produces a 480p pass at no credit cost with a short per-user cooldown, and the export charge lands once on the version you confirm regardless of clip count — the mechanics are in previews and the final export.