Spec-sheet claims that don't survive testing
Six spec numbers that promise more than they deliver: max duration, resolution, reference count, fps, native audio, multi-shot, and one test to run on each.
There are nine video models in the Versely catalog where the headline maximum resolution and the top resolution you can actually select at dispatch disagree with each other. Grok Imagine's video entry carries a 720p ceiling in one field and offers 1080p in the list you pick from. Seedance 2.0 carries 1080p in one and offers a 4K option in the other. Neither field is lying. They answer different questions, and only one of them binds the render you are about to pay for.
That is the shape of almost every spec-sheet disappointment. The number is real; the thing you assumed it guaranteed is not. Below are the six claims that cause the most rework, what each one actually asserts, and the single test that settles it before you build a shot list on top of it.
The two ceilings: duration and resolution
Max duration is the most over-read number on any model page. Across the 146 video models in the catalog, 122 declare a duration list and 24 declare none at all. The declared ceilings cluster hard: 41 models top out at 10 seconds and 43 at 15 seconds, with a thin tail at 20, 25 and 30.
Two things that ceiling does not tell you. First, it is a cap on one dispatch, not a promise that all of those seconds hold together. Second, it says nothing about granularity, which is what actually constrains an edit. Vidu Q3 and Wan 2.6 offer 5s, 10s and 15s and nothing between. Sora 2 steps in fours. Kling V3 and Happy Horse expose every whole second from 3 to 15. If your safe duration turns out to be 11 seconds, a coarse ladder makes you buy 15 and trim, or buy 10 and lose the beat.
The test: generate one clip at the declared ceiling with your real subject, then scrub it frame by frame and mark the first timestamp where identity, background layout or gait changes. That timestamp is your working ceiling. Compare it against the model's duration ladder to see whether you can actually land on it.
Output resolution has the opposite problem: too many fields. Fifty-three of the 146 video models declare no headline ceiling at all, and among the 92 that declare both a ceiling and a selectable quality list, nine genuinely disagree once you set aside models that simply spell 4K as 2160p. Sora 2's image-to-video entry is the disagreement that costs you: it carries a 1080p headline ceiling and offers exactly one selectable quality, 720p. The headline is the higher number and the selectable list is the one that binds. MiniMax H3 goes wrong a different way, declaring no headline ceiling at all while its selectable options run 768p, 2K and 4K, so "2K model" is true and still skips over a base tier well below 1080p.
The test: dispatch one clip at the top selectable tier and read the delivered file's pixel dimensions. Not the label, the dimensions.
The inputs: reference slots and native audio
Reference count is a capacity figure, not an influence budget. Nineteen video models in the catalog accept reference files, and the image caps run from 3 to 10: VEO 3.1 reference-to-video takes 3, Kling O3 Standard takes 4, Wan 2.7 takes 5 images and 5 video clips, and the Seedance 2.0 family takes 9 images plus 3 video and 3 audio references. A slot that exists is not a slot that changes the output.
The test: fix the prompt and the seed, then run the same shot with one reference, three references and the full complement. Compare identity on the subject you care about. If frames two and three look the same as frame one, the extra slots are costing you upload time and nothing else.
Native audio is a boolean stretched over a spectrum. Sixty-seven of the 146 video models carry an audio flag, 18 explicitly carry none, and 61 declare nothing either way. A flag that reads true tells you a track may come back. It does not tell you the dialogue will be in the language you prompted, that lip movement will land on syllables, or that the specific sound you named will be in the mix rather than generic room tone.
The test: prompt one line of dialogue and one named diegetic sound in the same clip. Check three things: whether the line is in the right language, whether the mouth matches at the frame level, and whether the named sound is audible and separate from the bed. Native audio that fails the third check is atmosphere, not sound design, and you will be adding a separate pass anyway.
The two nobody checks: frame rate and multi-shot
Frame rate is declared by 48 of the 146 video models, and the declared values are 24, 30, 48 and 60. That list is a menu of requests, not a guarantee about the file. Versely's default video frame rate is 25 fps, which is not on any model's declared menu, and that mismatch is exactly why the declared value is the wrong thing to plan around. What matters is the frame rate of the file that lands and whether it survives conforming to your timeline.
The test: probe the delivered file for its actual frame rate, then play a lateral pan back at full speed and look for judder. A clip that reads 24 in a field and delivers something else will look fine in isolation and stutter the moment it sits next to other footage.
Multi-shot is the claim with the least infrastructure behind it. Exactly one entry in the entire catalog carries a multi-scene feature tag, and it is Sora 2's storyboard mode, which is a distinct input format rather than a general "this model will cut for you" capability. When a launch post says a model handles multiple shots, it usually means the model will sometimes produce a cut when you ask for one, at a rate you have to discover yourself.
The test: write a prompt containing one explicit cut ("wide shot of the kitchen, cut to close-up of hands on the counter") and run it five times on the same seed family. Count how many outputs contain exactly one cut, how many contain zero, and how many contain three. That ratio is the real capability.
One test per claim
| Claim on the sheet | What it actually asserts | The test | You pass when |
|---|---|---|---|
| Max duration | A dispatch cap, not usable seconds | Scrub to first drift at the ceiling | Drift timestamp lands on a selectable duration |
| Output resolution | A label, sometimes across two fields | Read the delivered file's dimensions | Dimensions match the tier you paid for |
| Reference count | Slots available, not influence applied | 1 vs 3 vs max at fixed prompt and seed | Identity visibly improves past three |
| Frame rate | A request the pipeline may reshape | Probe the file, watch a lateral pan | No judder after conforming |
| Native audio | A track may return | Dialogue line plus one named sound | Right language, frame-accurate lips, named sound audible |
| Multi-shot | Cuts sometimes happen | Five runs, one explicit cut in the prompt | Cut count is consistent across runs |
What the six answers are for
Six tests is roughly one working session per model, and the output is not an opinion. It is six numbers you write down next to the model name: working duration, delivered resolution, useful reference count, delivered fps, audio verdict, cut reliability. That row is what you hand to whoever builds the shot list, and it is what stops the same model getting re-litigated every time someone reads a launch post.
Run the tests in one batch rather than one at a time. The agent will take a single prompt and fan it across several named models in one request, which is the cheapest way to get comparable outputs when the whole point is that the variable under test has to be the model. The 30-minute evaluation checklist is the version of this to run when a brand-new model lands and you need an adopt-or-skip call the same day.
FAQ
If the spec fields disagree, which one should I trust?
The one you can select at dispatch. A headline maximum is usually the highest number the model family has ever been documented at, across every mode and tier; the selectable quality list is the set of options the dispatch will accept. When they conflict, the selectable list is the one that determines what renders. Then verify with the delivered file, because that is the only field that cannot be wrong.
Is a 15-second ceiling better than a 10-second one?
Only if the extra five seconds are usable, and that is a property of the shot rather than the model. A locked-off product shot on a plain background often holds for the full ceiling. A walking subject with a busy background frequently does not. Test your actual shot type at the ceiling before you treat the higher number as an advantage, and check the duration ladder too: a model that offers 5s, 10s and 15s and nothing between is worse for editing than a lower ceiling with whole-second granularity.
Do more reference slots always mean better character consistency?
No, and the diminishing-returns point arrives earlier than the slot count suggests. Extra references help most when each one adds genuinely new information: a different angle, a different lighting condition, a detail the others do not show. Nine near-identical frames of the same pose mostly add upload time. Run the 1 vs 3 vs max comparison on your own subject before committing a campaign to a high-slot model.
How often do these answers go stale?
Whenever the model version changes, and quietly whenever a provider ships a serving change under the same name. Re-running six tests on your top three models once a quarter is cheap insurance. The answers that move most are duration drift and audio behaviour; resolution and reference caps tend to hold until the version number does.
The catalog's duration and resolution indexes show which models qualify for your constraints. Everything after that is a test, not a read.