Designing a paid trial task for AI creatives
A portfolio proves little when one striking frame is an afternoon's work. A trial task testing selection, iteration discipline and brand adherence under a cap.
Any candidate with credits and an afternoon can produce one arresting frame. That is no longer a filter, and a portfolio built from arresting frames tells you almost nothing about how someone behaves on a Tuesday with a brand manual, a deadline and a brief that will not cooperate.
A trial task fixes that, but only if it is designed to test the things that stayed scarce. Most trials still test production — make us something — which measures the one capability the tooling already provides. The trial worth running measures selection, iteration discipline and brand adherence, and it uses a fixed attempt budget as the instrument for all three.
What the trial has to measure
Three behaviours, in this order of predictive value.
Selection. Given more candidates than they can ship, can they pick correctly against a written standard and say why? This is the highest-signal behaviour and the least visible in a portfolio.
Iteration discipline. When something fails, do they diagnose and change one variable, or do they reroll and hope? This maps directly onto your credit spend and your calendar, and it separates cheap operators from expensive ones producing the same output.
Brand adherence under temptation. Will they ship the correct thing over the striking thing? The specific failure mode of generation-heavy work is beautiful off-brief output defended after the fact.
Everything else — prompt phrasing, model preference, software fluency — is ramp-time detail. Do not build the trial around it.
The task spec
Give them a real brief for work you have already done, so you have a known-good comparison and no live client is exposed. Four hours of their time is enough. More is imposition, less is not enough to see iteration behaviour.
- Supply a brand manual and a spec sheet. Colour, type, logo treatment, tone, forbidden imagery, aspect ratios, duration, platform. If your written standards are thin, this exercise exposes that first, which is useful information about you rather than them. A configured brand kit makes the constraint concrete rather than aspirational.
- Set a fixed attempt budget in credits. Not unlimited, not generous. Somewhere near what a competent operator would need with a little room, which you can size from your own logs by estimating the credit cost before dispatching a batch. The budget is the instrument. It is what makes the trial informative.
- Require three deliverables, not one. The finished asset; a shortlist of two or three near-misses with a line on why each lost; and a short log of what they tried and abandoned. The log is the part you will actually learn from.
- State one constraint that is genuinely awkward. A product that must stay consistent across shots, a mandated aspect ratio the model handles badly, a duration that forces a hard cut. You are not testing whether they can execute the easy version.
- Give them a named review point at the halfway mark. Fifteen minutes, optional. Whether they use it, and what they bring, is data. Candidates who arrive with a specific question about the brief are showing you how they will work.
- Do not specify models. Model choice under constraint is part of the assessment. If they ask, the answer is "whichever gets you there inside the budget, and say why in the log."
Everything interesting in this trial comes from the cap in step two. Unlimited attempts turn the exercise into a test of persistence, which is the thing generation made cheap and therefore the thing you do not need to measure.
Community-reported figures put usable rates near one clip in four, with three to five attempts per usable clip and more for character continuity or complex motion. Those are aggregated anecdote rather than measured data, so do not set your budget from them. Set it from your own logs, or from running the brief yourself first.
Two structural points.
The budget should bind, but not by much. A cap a competent person clears comfortably measures nothing; a cap nobody can clear measures your brief. You want the band where sequencing starts to matter, where twelve attempts on the hero shot means accepting a weaker second one.
Tell them the budget is real and will not be extended. Half the signal is allocation. A candidate who spends the first third establishing which model handles the awkward constraint, then executes efficiently, is showing you how they will work on a live account.
One practical note if your trial involves an edit: the editor's 480p preview pass is free to run and carries a short per-user cooldown, and the final export is charged once regardless of clip count. That means preview iteration does not eat the attempt budget, which is worth stating explicitly in the brief so candidates do not ration the wrong thing. The mechanics are in previewing before you pay for the export.
Scoring rubric
Score before you discuss, and score independently if more than one person reviews. Five criteria, weighted.
| Criterion | Weight | What a strong result looks like |
|---|---|---|
| Brief adherence | 30% | Every spec-sheet requirement met, deviations flagged in the log with reasoning |
| Selection quality | 25% | The shortlist reasoning names specific defects; the chosen asset is defensible against the near-misses |
| Iteration discipline | 20% | Log shows single-variable changes after failures, not blind rerolls; budget spent unevenly and deliberately |
| Craft | 15% | Composition, pacing, legibility at delivery size hold up next to your known-good version |
| Communication | 10% | Log is readable; questions asked were the right ones; assumptions stated rather than buried |
Craft is deliberately not the largest weight. It is the criterion most correlated with a candidate's previous access to good tooling and least correlated with how they will perform under your constraints. It still matters — someone who cannot see a bad cut has no standard — but weighting it top turns the trial back into a portfolio review.
Run the finished asset through your normal brand-manual audit rather than judging it by eye. If the candidate's work passes a check your own output sometimes fails, that is worth knowing.
Paying for it, and the paperwork
Pay for the trial. Not only for fairness; it changes what you can ask for and how seriously the work is taken. There is no reliable published day-rate benchmark for AI-first creative production specifically — the widely quoted hourly figures in this space are consulting rates, not production rates — so anchor on what you pay your own people for equivalent hours rather than on a number from a pricing blog.
Three pieces of paperwork are worth having in place before the brief goes out.
Ownership, stated accurately. Do not promise or demand more than exists. Model providers assign contractual ownership of outputs to the user under their terms, but a contract cannot create copyright that statute does not grant, and the US Copyright Office position is that protection attaches to human contribution — meaningful modification, human-authored elements, and the creative selection and arrangement of generated material. Write the trial agreement so it assigns whatever rights exist without warranting that the raw generated portions are copyrightable. The general shape of this is covered in legal and licensing for AI content.
Usage scope. State plainly whether trial output may be used commercially. If it may, pay accordingly. If it may not, say so in writing, because a candidate who later sees their trial work in a live campaign will tell people. The vocabulary for this sits under usage rights.
Provenance handling. Ask them to deliver with embedded credentials intact and note any step in their process that strips them. This doubles as a small competence test and it costs them nothing.
Five ways trial design goes wrong
- Testing on your hardest live problem. Tempting, and it turns the trial into free consulting. Use finished work.
- No known-good comparison. Without your own version of the same brief, scoring drifts toward whoever produced the prettiest thing.
- Unlimited attempts. Measures stamina, not judgment.
- One reviewer scoring after discussion. Anchoring destroys the rubric. Score first, discuss second.
- Asking for volume. "Give us ten variations" rewards the behaviour you are trying to screen out. Ask for one shipped asset and the reasoning behind the ones that lost.
Take delivery as links rather than archives, which also gives you a clean record of what was submitted when. Shareable review links do this without asking a candidate to create accounts on your stack.
FAQ
How long should the trial be?
Four hours of candidate time, delivered inside a week. Long enough to see allocation decisions across more than one shot, short enough that people with jobs can do it. Past a day, you are selecting for availability rather than skill.
What if a candidate refuses a trial task?
Reasonable people refuse unpaid ones. Refusal of a paid, scoped, four-hour trial is itself information, though not always negative — senior candidates with strong references sometimes have enough evidence in their track record. In that case substitute the shortest version: hand them twenty assets and written criteria and score the reasoning behind their cuts.
Should we let them use their own tools?
Yes, and note what they choose. Forcing your stack tests ramp speed rather than judgment. Just require the log to name what was used, so the result is reproducible.
How do we compare candidates who chose different models?
Score the outcome and the reasoning, not the model. Two candidates hitting the same brief inside the same budget with different routing are both correct. What differentiates them is whether the log shows they chose deliberately, and whether their reroll-or-fix decisions were diagnostic. Model choice is a means; budget-adjusted outcome is the score.
Build the trial from a brief you have already shipped, set the cap from your own logs, and score the log as carefully as the asset. The finished piece tells you what they can make. The log tells you what they will cost.