Strategy

    Charging for model research as a line item

    Finding which model holds a brief is billable work most studios absorb. How to scope a model-selection pass, cap it on three axes, and invoice it defensibly.

    Versely Team9 min read

    Before a single deliverable gets made, someone spends half a day finding out which model can hold the brief. That half day is real work with a real output, and almost nobody invoices it. It disappears into "setup," or into the first deliverable's margin, or into a Sunday.

    It is also the work that most determines whether the project is profitable. Route a campaign to a model that needs eight attempts per usable shot when a different one needs three, and you have committed to a cost structure for the whole engagement in an afternoon nobody was paid for.

    The catalog side of this is not small. Versely lists over 300 model pages, and the useful shortlist for any given brief is rarely obvious from a leaderboard. Boards measure aggregate preference on generic prompts. Your brief is not a generic prompt.

    What a model-selection pass actually contains

    Write it out once and the case for billing it makes itself:

    • Translating the brief into three to five representative test prompts, including the hardest shot rather than the easiest
    • Building a shortlist from the catalog based on what the brief needs — duration, audio, reference conditioning, aspect ratio, motion control
    • Dispatching the suite across the shortlist
    • Scoring outputs against a rubric rather than by vibe
    • Measuring attempts-to-usable per model, which is the number that actually decides cost
    • Checking licence and usage terms against how the client intends to distribute
    • Writing a recommendation someone else could act on

    That is a day's work with a professional output. Presented as a line item it looks like exactly what it is. Absorbed into a project fee it looks like nothing, which is why it is the first thing to get squeezed when a deadline moves.

    Scope it before you price it

    Three artifacts turn an open-ended research task into a bounded one.

    The shortlist. Filter the catalog by hard requirements first, not by reputation. Does the brief need native audio? Reference-image conditioning? A specific aspect ratio? Duration beyond a short clip? Each of those cuts the field hard, and filtering by capability before quality is what keeps the pass finite. The model catalog and the comparison pages are the starting filter; category pages like the best text-to-video models narrow further, though they answer a general question rather than yours.

    The prompt suite. Fixed, written before any generation happens, and reused across every model so the comparison means something. Include the shot you are least confident about — a suite made of easy shots ranks models on things that were never going to fail. A twenty-prompt suite for testing video models is the structured version; for a single client pass, three to five prompts drawn from the actual shot list is usually enough.

    The rubric. Score the same attributes for every output, on a fixed scale, before you look at the next model. Prompt adherence, subject fidelity, motion plausibility, artefact rate, and attempts-to-first-acceptable-take. A scoring rubric for comparing video models is the version worth adopting wholesale, because a rubric you invent mid-test drifts toward whatever you have already decided.

    All three exist before the pass starts. That is what makes it a scope rather than a research budget.

    Cap it on three axes and a clock

    Model research is the easiest work in this business to do forever. Cap it explicitly:

    1. Models. A named shortlist, usually four to six. Not "we'll test what looks promising."
    2. Prompts. The fixed suite. Adding prompts mid-pass invalidates the comparison anyway.
    3. Attempts per prompt per model. Three is a reasonable default. It is enough to distinguish a bad model from an unlucky roll, and few enough to keep the credit spend bounded.

    Four models times four prompts times three attempts is 48 generations. That is a number you can forecast in credits before you quote, which matters because every generation on Versely costs credits and there is no free allowance sitting behind the pass. Then put a clock on it: two working days, findings delivered on the second.

    The cap has a second function beyond cost control. It forces the pass to be conclusive. An uncapped comparison keeps finding new candidates and never produces a recommendation, which is the failure mode that makes clients suspicious of research line items in the first place.

    The single biggest efficiency here is dispatch. The agent chat can fan one prompt across several named models in a single request, which turns the slowest part of the pass — running the same suite four separate times — into one operation. Agent routing or picking the model yourself covers when to hand the choice over and when to keep it, and A/B testing models on one prompt is the mechanics.

    Three ways to invoice it

    Structure How it works Best when
    Standalone line item Fixed fee, its own deliverable, invoiced on delivery of the memo New client, unfamiliar brief, or a client who is choosing between vendors
    Credited against production Charged upfront, deducted from the production fee if the project proceeds Sales-sensitive situations; removes the "paying to be quoted" objection
    Folded into a setup fee One line inside onboarding, alongside style lock and asset library Retainers, where the pass is genuinely one-time

    The credited-against structure is the one that closes deals. The client hears that the research is effectively free if they proceed, and you get paid either way. It also fixes an incentive: a vendor who is paid to select a model is being paid to be right, whereas a vendor absorbing the pass has a quiet incentive to route to whatever they used last week.

    For ongoing relationships, the setup-fee route is cleanest. Month one of a retainer is mostly unbillable work of exactly this kind, and charging a setup fee before the first retainer month is the general case that model research fits inside.

    One structure to avoid: hourly. Research is precisely the work where getting faster should not cut your fee, and a client watching an hourly meter on exploratory work will start asking why it took four hours instead of two.

    The memo is what makes it defensible

    A line item with no artifact behind it feels like padding. Produce a written selection memo and it stops feeling like anything — it is a document the client can file, forward and act on.

    One page:

    1. The requirement, in the brief's own terms. What the shots need, not what the models do.
    2. The shortlist and why each model made it. One line each. Naming what you excluded is often more persuasive than naming what you tested.
    3. The scores, as a table. Same rubric, same prompts, same attempt count for every row.
    4. Attempts-to-usable per model. The cost-relevant number, and the one nobody else will give them.
    5. The recommendation, with the runner-up and the condition under which you would switch.
    6. Constraints found. Duration limits, aspect-ratio behaviour, audio handling, anything in the terms that affects how the client can distribute the output.

    Point 4 is the one that justifies the fee on its own. Published boards measure preference between outputs; they do not measure how many attempts it took to get an output worth preferring, and that ratio is what determines project cost. Usable rate is the benchmark nobody publishes is the longer argument for why your own logs beat any public leaderboard here.

    The memo also has a private value. Every pass you run and file is a shortlist you do not have to rebuild next time a similar brief lands, which is what turns a billable service into an asset.

    When not to charge for it

    Three cases where the line item is the wrong move:

    You already know. If the brief is a near-copy of work you did last month, routing is not research. Charging for a decision you made in ninety seconds is the kind of thing that gets noticed on the third invoice.

    The pass is one model deep. Testing your default against nothing is not a selection pass. Either run a real comparison or do not bill for one.

    The client is buying an outcome, not a process. Some clients want a single number and no visible machinery. Fold the research into the deliverable price and skip the itemisation — the point is to be paid for the work, not to win an argument about how invoices should be formatted.

    The rest of the time, itemise it. Model research is the most leveraged hour in the project and the one most consistently given away.

    FAQ

    How often does a selection pass need redoing?

    Per brief when the brief is genuinely different, and roughly quarterly for a standing client whose work is consistent. Model capability moves fast enough that a routing decision from six months ago deserves rechecking, and a small re-test against your existing rubric is much cheaper than the original pass because the suite and the scoring already exist.

    Won't clients ask why they are paying me to learn my own tools?

    The framing to avoid is "learning." You are not learning the tools; you are testing candidates against their specific brief, which is work that could not have been done in advance because their brief did not exist. The memo makes that obvious. If someone still objects, the credited-against-production structure removes the objection entirely.

    What if the recommendation turns out wrong in production?

    Say so early, in writing, and route to the runner-up you already named. This is exactly why the memo lists a runner-up and a switching condition. A pass that produces one answer and no fallback is a pass that cannot survive contact with a shot list, and building the fallback in at the start costs nothing.

    Should the memo name the credit cost per model?

    Give the client the comparison in terms they can act on: relative spend and attempts-to-usable per model, plus the total credit budget the pass consumed. That is enough to explain why the recommended model is the cheaper route without turning the memo into a spreadsheet the client will try to renegotiate against.