Workflows

    WIP limits for a generation-heavy studio

    Cheap generation invites starting more jobs than anyone can review, so everything lands late together. Setting WIP caps per review stage, and holding them.

    Versely Team9 min read

    The symptom is specific enough to diagnose from across the room: nothing is late individually, and then eleven things are late in the same week. Every job looked healthy at standup because every job was moving. They were all moving through a stage that can process four things at once while holding nineteen.

    That pattern arrived with cheap generation. When starting a job cost real money and a week of someone's time, the studio rationed starts by accident. Now a start costs a prompt, so nothing rations them, and the number of jobs in flight drifts upward until it settles at whatever the team can tolerate rather than whatever it can finish.

    Why cheap starts break flow

    The relationship is arithmetic, not cultural. Little's Law holds that the average number of items in a system equals throughput multiplied by the average time each item spends in it. Rearranged for the thing you care about: average cycle time equals work in progress divided by throughput.

    Throughput here is set by your slowest stage, which in a generation-heavy studio is review, not production. So if review can finish six assets a day and you hold sixty in flight, average cycle time is ten days. Add twenty more starts and cycle time goes to thirteen days without a single asset being produced any slower. You did not add capacity by starting more work. You added waiting.

    Two things make this worse than the classic manufacturing version.

    Generated work decays while it queues. An asset sitting in review for nine days was briefed against a campaign that has since shifted, and a proportion of the queue gets re-briefed rather than shipped. That is throughput spent on nothing.

    Context switching is the real review cost. Reviewing forty assets across eleven campaigns is slower per asset than forty across three, because the reviewer reloads brand context, brief and history each time. High WIP does not just delay review, it makes review less efficient, which lowers throughput, which raises cycle time again.

    Cap the review stages, not the generation stage

    The instinct is to limit generation. Wrong lever. Generation is not the constrained stage, so capping it does not raise throughput, and exploratory generation is genuinely useful — batch-testing twenty creative directions to find two worth pursuing is a good use of credits, and the batch-testing pattern depends on producing widely.

    The distinction that makes this work: cap the number of jobs that have entered a stage requiring human attention. Generating fifty candidates for one job is one item of WIP. Generating five candidates each for ten jobs is ten. The first is exploration; the second is ten commitments you have to see through.

    So the unit is the job, and the caps live where a human has to look and decide.

    Setting the numbers

    Do not pick round numbers. Derive them, in four steps per stage.

    1. Measure review capacity in items per day. Reviewer hours actually available for review, divided by minutes per item. Be honest about hours available. A creative director with four meetings a day has perhaps ninety minutes of real review time, not six hours.
    2. Decide the cycle time you want in that stage. Usually one to two days. Longer than two and the asset is stale relative to the brief; shorter and you are likely to be under-reviewing.
    3. Multiply. Cap equals capacity per day times target days in stage. Six items a day at a two-day target is a cap of twelve.
    4. Round down. The formula gives the point at which the stage is exactly saturated, and a saturated stage cannot absorb variance. Take roughly 80% of it.

    A worked example for a three-person team, all numbers illustrative:

    Stage Capacity/day Target days Raw cap Working cap
    Shortlist (curation) 10 jobs 1 10 8
    Internal creative review 6 jobs 1.5 9 7
    Client or stakeholder approval 4 jobs 2 8 6
    Delivery and provenance check 12 jobs 0.5 6 5

    Total WIP across the board is 26 jobs, which will feel low to a team currently holding sixty. It is not low. It is what they were finishing anyway, minus the queue they were carrying.

    The delivery cap catches people out. It looks administrative, so it gets no limit, and then launch week arrives and every asset needs metadata and disclosure verified at once. Provenance checks are fast per item and unforgiving in aggregate.

    Which stages need their own cap

    Give a cap to any stage where an item can sit waiting for a person. In practice that is four, sometimes five.

    Shortlist. Between generation and creative review. This is the stage most studios do not name at all, which is why candidate sets pile up unfiltered. The cap here is what stops sixty candidates arriving on a director's desk.

    Creative review. Internal sign-off against the brand standard. Serial routing kills this stage; parallel review with one named decider per asset class is the fix, and the mechanics are in approval workflows that do not stall.

    Client or stakeholder approval. The stage you control least and the one that most needs an explicit limit. Items wait on someone outside the studio, and without a cap the queue grows to whatever the client's inbox absorbs.

    Delivery and provenance. Metadata, disclosure labels, platform specs, naming. Fast per item, dangerous in batches. Pair the cap with asset naming and version discipline so the checks are mechanical rather than investigative.

    Rework, if it exceeds a fifth of your volume. Reworked assets behave differently from first-pass ones and otherwise consume review capacity invisibly. Give them their own lane and limit.

    Generation does have a natural limit that is not a WIP cap: provider concurrency. That is a scheduling constraint rather than a flow constraint, and it belongs in model selection rather than here.

    Holding the cap when it is hit

    This is where the practice succeeds or gets abandoned in a fortnight. Hitting the cap must produce a defined behaviour, not a discussion.

    Nothing new enters. The team helps clear the blocked stage instead. That is the entire mechanism. Someone whose next job cannot start goes and reviews, or curates, or runs delivery checks. It feels wrong the first three times and is the point of the system.

    Log every breach with its reason. A stage that hits its cap twice a week has a capacity problem; a stage that hits it twice a quarter is correctly sized. Three months of breach logs will tell you where to hire more reliably than any headcount ratio.

    Do not raise the cap to relieve pressure. Raising it lengthens cycle time and reduces per-item attention. The only legitimate reasons to raise a cap are more reviewer hours or faster per-item review.

    One budget-side effect: capping jobs in flight makes credit spend predictable, because dispatch is bounded by what can be reviewed. If you already estimate credit cost before dispatching a batch, the cap turns that estimate into a forecast rather than a lower bound.

    The cap will also be tested by a client who wants one more thing started this week. Three responses that work, in order of preference.

    Trade, do not add. "We can start it today if we pause the product refresh. Which first?" This usually resolves the conversation, because the client is choosing between real options rather than being refused.

    Show the queue. Most clients pushing for a new start believe capacity is idle. A one-screen view of each stage changes the conversation from willingness to arithmetic. A defined revision policy helps here too, since much surprise WIP is revisions never scoped as new work.

    Quote the queue time. If neither works, accept the job and state the honest date given current WIP. "We can start it now and it will land the 14th, or start it Monday and it lands the 9th" is counterintuitive and true, and it is a better answer than a date you will miss.

    What does not work: quietly accepting and hoping. That converts an explicit conversation into eleven things late in the same week, which is exactly the failure you set the caps to prevent.

    FAQ

    Is this just Kanban?

    The mechanism is, yes. What changes in a generation-heavy studio is where the limits belong. Classic Kanban limits the whole board fairly evenly because every stage costs roughly comparable effort to enter. Here, one stage got dramatically faster and the rest did not, so the caps cluster downstream and that stage stays deliberately uncapped.

    What if we are one person?

    The caps matter more, not less, because there is no one to absorb overflow. Two limits are usually enough: jobs awaiting review and jobs awaiting client sign-off, both set at two or three. The solo failure mode is starting six things on a good day and having six half-finished on a bad one.

    How do we count a job with twelve variants?

    One job through review if it is reviewed as a set against one brief, twelve if each needs an independent decision. Count by decisions required, not by files. A single ad concept in six aspect ratios is one item; six distinct concepts are six.

    Will this reduce our output?

    It will reduce your starts and should not reduce your finishes, because finishes were already governed by review throughput. What usually changes is that cycle time drops sharply, rework falls as reviewers give each item more attention, and the ratio of approved to generated assets improves. If total finishes actually fall, the cap was set below real capacity — recheck step one. Your first-pass usable rate is the other number to watch alongside it.

    Count what is in flight right now, then count what you finished last week. If the first number is more than about three times the second, you are not busy. You are carrying a queue.