Tier deliverables by difficulty, not duration
Runtime stopped predicting effort. Build a rate card on the nine shot attributes that actually drive attempt counts, with tier bands and an attempt allowance.
A 10-second shot of a hand pouring liquid into a glass can take twenty attempts. A 30-second drone push over a coastline can land on the second. Price both by the second and you have built a rate card that charges triple for the easy one.
Duration was a reasonable proxy for effort when a camera had to be somewhere for a length of time. In a generated pipeline the relationship is close to gone. Runtime now predicts almost nothing except render time, and render time is not where the work lives. The work lives in attempts, and attempts are driven by what is in the shot.
What actually drives attempt count
Nobody has published measured hit rates for generated video. What exists is community reporting — figures in the range of one usable clip in four, and three to five attempts per usable shot, with worse ratios on specific character action and complex motion. Treat those as anecdote, not benchmark. The reason to take them seriously anyway is that they are directionally consistent with what anyone running volume observes: the spread between easy and hard shots is enormous, and it tracks attributes, not seconds.
Nine attributes carry most of the variance:
- Human faces held in frame, especially in motion. The failure threshold for a face is far lower than for anything else on screen.
- Speech with visible mouth movement. Adds a whole failure class on top of the face.
- Hands doing something specific. Grip, count, articulation. The single most reliable way to burn attempts.
- Physical contact and object interaction. Weight, gravity and contact points are where generated motion still gives itself away — the failure modes are catalogued in physics failures in generated motion.
- A real SKU that has to stay itself. A packaged product with a fixed shape, label and colour is a hard constraint the model does not know it is under. Keeping a physical SKU believable is a job on its own.
- Legible on-screen text. Sometimes fine, frequently not, and never reliable enough to promise — see what holds up in generated on-screen text.
- Character consistency across shots. One shot of a person is easy. The same person in six shots is a different deliverable.
- Continuity with an adjacent shot. Matching lighting, wardrobe and lens between two separately generated clips is work that has no equivalent in the single-shot case.
- Named brand assets. Logos, typefaces, packaging artwork, anything where "close enough" is a legal problem rather than a taste problem.
Notice what is missing: duration, resolution, aspect ratio. Those change credit spend, which matters to your input cost, but they barely move the number of attempts.
The tier bands
Score each shot by counting how many of the nine attributes it carries, then band it. The exact multipliers are yours to set against your market; the structure is what transfers.
| Tier | Typical shot | Attributes present | Attempt allowance | Multiplier on base |
|---|---|---|---|---|
| T1 Ambient | Landscape, texture, abstract motion, atmospheric b-roll | 0 | 2 | 1x |
| T2 Controlled subject | One person, simple action, no speech, generic wardrobe | 1–2 | 4 | 1.5x |
| T3 Interaction | Hands on a product, two subjects, contact motion | 3–4 | 8 | 2.5x |
| T4 Hero | Real SKU, legible pack copy, brand-critical framing | 5–6 | 12 | 4x |
| T5 Bespoke | Recurring character across a sequence, matched continuity, dialogue | 7+ | Quoted per shot | Quoted |
T5 is deliberately not a multiplier. A shot with seven of these attributes is not a harder version of a normal shot; it is a small project with its own test phase. Quoting it off a card is how you end up eating a fortnight.
The attempt allowance is the part most rate cards leave out, and it is what makes the tier mean something. It converts a vague "harder" into a specific number of generations included in the fee, which gives you a defensible boundary when a shot goes past it. Past the allowance, you are either in a change order or in a fallback — never in unpaid overtime.
Auditing a shot list at quote time
The tier is only useful if it gets assigned before the quote, not after the shot fails. Fifteen minutes with the shot list does it:
- Read every shot and mark the attributes. Nine checkboxes, one line per shot. Do not estimate difficulty holistically; count.
- Circle the SKU shots. Any shot where a real product must look like itself is at minimum T4 regardless of what else it carries.
- Look for repeats of the same subject. Three shots of the same character is not three T2s. It is a continuity requirement, which moves all three up a tier.
- Flag anything with text in it. Then decide, at quote time, whether the text is generated or composited in the edit. Compositing it later usually drops the shot a full tier.
- Total the tiers, not the seconds. The output is a mix — for example 4×T1, 6×T2, 3×T3, 1×T4 — and that mix is the quote.
Step 4 is where most of the money is. A large share of shots that look like T4 are only T4 because someone decided the pack copy had to be generated rather than laid over in the timeline. Versely's editor is EDL-based, so text and graphics live on the timeline and are re-renderable without regenerating the underlying clip. Moving text out of the generation is a pricing decision as much as a craft one.
What the tiers do to the rest of the quote
Once difficulty is the unit, three other clauses get easier to write.
Revisions attach to tiers. A T1 revision is a reroll. A T4 revision may be a reshoot of the whole setup. One flat revision policy across both is a policy that is too generous on one end and too strict on the other. Band the revision allowance the same way you banded the rate, then fold it into your revision policy so there is one document, not two.
Substitutions become tradeable. When a client wants to add a shot mid-project, a tiered card gives you an exchange rate. Two T1s out, one T3 in. That is a conversation about creative priorities rather than about your fee.
Credit budgeting stops being guesswork. Attempt allowance times model spend per attempt gives you a forecastable input cost per tier, and Versely bills every generation in credits with no free allowance, so the forecast is real money. Reroll rates is the method for measuring your own attempt multiplier rather than borrowing someone else's, and estimating credit cost before you dispatch a batch is the pre-quote check. The worked scenario for a finished 30-second ad gives you a floor.
Keep your own attribute weights
The nine attributes above are the starting set, not a law. Your weights will drift from anyone else's, because attempt counts depend on which models you route to, how you prompt, and what your clients ask for.
The fix is to log one number per shot: attempts to first acceptable take. Nothing else. Over a couple of months that column tells you which attributes are actually expensive in your pipeline, and some of them will surprise you — a model that struggles with hands may be fine with reflective packaging, and the reverse also happens. That is the same discipline behind usable rate as a benchmark, applied to pricing instead of model selection.
Rebase the tiers quarterly. Model capability moves, and an attribute that was T4 in March can be T2 by September. When that happens, the correct response is to move the shot down a tier and keep the rate card's overall shape, not to cut your prices — the client is buying a finished shot, and the difficulty of producing it was always your business rather than theirs.
FAQ
Do I show the client the tier table?
Show the tiers and the attempt allowances. Do not show the attribute checklist. The tiers explain why one shot costs more than another, which is useful and builds trust. The checklist invites a line-by-line negotiation about whether a hand is really in frame, which is nine arguments per shot and helps nobody.
How do I handle a client who only wants one flat per-video price?
Give them one, priced off the tier mix of a representative video, and write a shot-composition assumption into the scope: so many T1, so many T2, so many T3 per deliverable. The flat price stands as long as the mix does. When the mix changes, the assumption is the trigger for a repricing conversation rather than a surprise.
Does duration matter at all?
For credits, yes — longer generations cost more, and clip length limits vary by model. For effort, only at the edges. A single continuous take past roughly eight seconds gets harder because the model has more time to drift, so treat long single takes as an attribute in their own right rather than as a linear price ramp. Versely's default timeline runs at 25 fps, which is worth knowing when you are matching generated clips against anything shot at another rate.
Where does audio fit in the tiers?
Separately. Speech with visible mouth movement is already an attribute, but voice, music and sound design are their own line and should be priced as such rather than absorbed into a video tier. Whether the model's native audio is good enough to skip a separate pass is a per-project decision, and it changes the finishing cost more than it changes the generation cost.