Hiring for taste, not tool skills
Tool fluency decays with every model release; judgment about what to keep does not. Interview questions and portfolio criteria that separate the two.
A candidate who is fluent in the video model you use today has a skill with a measurable expiry date. The catalog turns over, prompt conventions shift, the thing that needed a careful negative prompt last quarter handles it natively this quarter, and the specific fluency you hired for is worth less every month. Meanwhile the candidate who can look at twenty generated options and correctly pick the two worth finishing has a skill that has not depreciated once.
Most job specs still lead with the first thing. A survey of roughly 1,300 freelance designers run between February and April 2026 found 62% using AI often or sometimes, and demand data over the same period shows AI-related freelance skills up 109% year on year. Tool use is table stakes now. Screening for it mostly screens for recency.
What tool fluency is actually worth
Not nothing. It is worth about a fortnight of ramp time, and it correlates loosely with curiosity, which does matter. What it does not do is predict output quality, because the expensive part of generated work is not producing a candidate — it is knowing which candidate to keep and when to stop.
There is a second-order problem with screening on tools. It selects for people who have optimised for the current stack, and those people are systematically slower to abandon it. In a category where model rankings move monthly, attachment to a particular tool is closer to a liability than an asset. What you want is someone who reads the model catalog as a menu rather than an identity.
The durable version of the skill is narrower than "taste" as usually used. Be specific about it.
Taste, defined as something you can score
Taste in this context is three separable behaviours, and you can test each one.
Rejection under a standard. Can they say why something fails, in terms specific enough that someone else could act on it? "It feels off" is preference. "The product label goes illegible below thumbnail size and the third act has no motion" is a standard being applied.
Knowing when to stop. Every generated asset has a point past which more attempts stop improving it. People who cannot find that point burn budget and calendar in equal amounts. This is measurable — attempts to acceptance, on the same brief, across candidates.
Holding a brief under temptation. The specific failure of generation-heavy work is producing something beautiful that is not what was asked for, and then arguing for it. A candidate who ships off-brief work because the output was striking will do that on your accounts.
None of the three requires knowing any particular tool. All three show up in a portfolio if you ask for the right thing.
Portfolio criteria that separate selectors from generators
The standard portfolio is now nearly uninformative, because one striking frame is available to anyone with credits and an afternoon. Change what you ask for.
| Ask for | What you learn | The tell |
|---|---|---|
| The rejected set alongside the final | Whether they have criteria or just outcomes | They can name the specific defect in each reject |
| Three pieces made under a brand manual they did not write | Whether they can subordinate their instinct to someone else's system | The work looks like the brand, not like them |
| One project where the first direction was abandoned | Whether they can be wrong quickly | They can say what the evidence was, and it was not "the client hated it" |
| Attempt counts on a finished piece | Their relationship with the reroll button | A plausible, non-round number, and a reason for it |
| Something they shipped that they still dislike | Whether they distinguish "good" from "correct for the job" | The dislike is aesthetic, the defence is strategic |
The reject set is the highest-yield request and almost nobody asks for it. A generator has a folder of finished work. A selector has a folder of finished work and can immediately explain the eleven things beside it that did not make it. The reasoning is the portfolio. The reel is the artefact — worth reviewing on its own terms, and the conventions of a video freelancer's reel still apply, but it is no longer the evidence.
One caution: a high attempt count is not failure and a low one is not skill. The number is a conversation opener, not a score.
Seven interview questions
Each of these has a wrong answer that sounds impressive.
1. "Show me a batch of your own output and cut it to two. Talk while you do it." The core exercise. You are listening for criteria stated before the cut, not rationalisation after. Weak answer: silence, then two picks. Strong answer: naming what the brief needed, then eliminating against it.
2. "What's the last output you loved that you didn't ship?" Tests whether brief adherence beats attachment. If they have no such example, either they always ship what they love, which is a governance risk, or they have not worked under a real brand system.
3. "A model you rely on gets deprecated tomorrow. Walk me through your first day." Tests portability of process. Weak answer: names a replacement tool immediately. Strong answer: describes re-establishing a baseline — same test prompts, same shot types, compare against known-good output. That instinct is the same one behind a scoring rubric for comparing video models.
4. "Tell me about a time you kept generating past the point where you should have stopped." Tests self-knowledge about the stop condition. Everyone has done it, so a candidate who claims otherwise is either new or not paying attention.
5. "Here's a brand manual and here's an asset. Would you ship it?" Give them a genuinely borderline asset. You are testing whether they check against a written standard or against their eye. The best answers do both, and separate them out loud: this is compliant and I would still push back for this reason.
6. "How do you decide between fixing a generation and rerolling it?" The most operationally predictive question on the list. It maps directly onto budget. A candidate with a real answer describes reading the specific failure and changing one variable, which is the reroll-or-fix judgment in practice. A candidate without one describes rerolling until it works.
7. "What would you have to see to change your mind about a creative direction you were confident in?" Tests whether their confidence is falsifiable. Anyone who cannot answer has opinions rather than judgment, and opinions do not update when the ad data comes back.
And one question for the reference check, asked of a previous manager: "When they disagreed with a creative decision, what did that look like?" You are listening for a specific shape. Disagreement that arrived early, in terms of the brief, and stopped once decided is what you want. Disagreement that arrived after delivery, in terms of personal preference, and continued afterwards is the profile that makes review cycles expensive — and expensive review is what a generation-heavy studio can least afford.
What to stop screening for
- Named-tool checklists in the job spec. They shrink your pool toward people whose skill is most exposed to obsolescence. State the problem shapes you work on instead.
- Prompt-writing tests as the main exercise. Prompt craft is real and it is the fastest-decaying item on the list. Test it as a component of the selection exercise, not as the exercise.
- Volume claims. "I produced 400 assets last quarter" tells you about their tooling. Ask what proportion shipped, and whether they know. Anyone tracking a first-pass usable rate at all is ahead of most of the market, and the argument for measuring usable rate applies to people as much as to models.
- Speed as a standalone virtue. Fast generation with weak selection produces more work for whoever reviews it. The relevant speed is brief to approved, not brief to output.
- Portfolio polish in isolation. A polished reel now costs an afternoon. Judge the reasoning behind it.
None of this means ignore craft. Someone who cannot see a bad cut or a broken composition has no standard to apply, and taste without craft knowledge is just confidence. The point is the ordering: craft and judgment first, current tool fluency as a ramp-time consideration rather than a filter.
FAQ
How do you test taste without an unpaid work assignment?
The batch-cutting exercise in question one uses their own existing output, takes fifteen minutes, and costs the candidate nothing. It is the highest-signal, lowest-imposition test available. Anything requiring them to produce new work for you should be paid and scoped tightly.
Isn't "taste" just a way of hiring people like us?
It can be, and that is the real risk in this approach. Guard against it by writing the criteria down before the interview and scoring against them, rather than scoring against agreement. If a candidate's cuts differ from yours but the reasoning references the brief correctly, that is a pass. Requiring the same picks is how a team ends up with one opinion and creative fatigue across every campaign.
Does this apply to junior hires?
The behaviours do; the evidence base does not. A junior will not have a reject folder or attempt logs. Substitute a supplied batch and written criteria, and score the reasoning. What you are looking for at that level is whether their reasons reference anything external at all.
What if the whole team lacks this and we are hiring the first one?
Then write the criteria first, because you cannot interview for a standard you have not articulated. An hour spent turning your brand manual into a pass-fail checklist makes every subsequent interview sharper, and it makes the audit pass against the brand manual something a new hire can run in week one.
Rewrite the top of your next job spec. If the first three lines name tools, you are hiring for the part of the job that expires.