Strategy

    Milk a format until the numbers say stop

    Most working formats get retired on boredom, not data. Plot a decay curve per format and use a rolling median to tell one dud episode from a spent shape.

    Versely Team9 min read

    The retirement notice for a working format almost always arrives from the person making it, not from the chart. You have shot the same opening eleven times. You can recite the caption. The next episode feels like filling in a form. So you kill it, announce a "new direction," and watch the numbers fall off a shelf.

    Paddy Galloway's framing in his Colin & Samir sit-down is blunt: borrow a proven format, then milk it while it delivers instead of abandoning it to copycats. While implies a test. Most creators default to how they feel on a Tuesday.

    This is the test. One number per format, plotted against episode number, smoothed with a rolling median, with a written stop rule you commit to before you are bored.

    Boredom peaks before the format does

    The reason the gut is a bad instrument here is that the two curves run out of phase.

    Your fatigue is a function of how many times you have built it. It starts at episode one and climbs. The audience's response is a function of how many times they have seen it, and that curve does not start climbing until enough of them have seen it twice. The earliest episodes are structurally not the peak.

    So the creator's boredom is highest exactly where the format's returns are still climbing. Six episodes in, you have done the work six times and the audience has barely learned the shape. That is the single most common moment a working format gets shot.

    There is a second failure that looks identical from the inside. A format that has genuinely decayed still occupies the slot a working one would fill, and audience fatigue is usually a format problem rather than a model problem — the fix is a different shape, not a better render. Only a number tells you which error you are making.

    Pick one number per format, before episode one

    The metric has to be chosen in advance, because choosing it afterwards means choosing whichever number confirms what you already decided.

    Pick the one that reflects whether people are still watching the shape, not whether one episode got lucky with distribution. Views are the worst candidate: they measure the feed lottery more than the format.

    Surface The number to track What the band looks like
    YouTube Shorts Viewed vs. Swiped Away, in the Shorts Feed tab of Studio The band most often quoted from Galloway's large-scale Shorts analysis is roughly 70–90%, with collapse under 60%. This is a practitioner band, not a published YouTube threshold
    TikTok Completion rate Analytics vendors commonly publish floors that scale with length: above ~50% average watch under 30s, ~40% for 30–60s, ~30% past 60s
    Instagram Reels Sends per reach Adam Mosseri has publicly named watch time, sends per reach and likes per reach as the ranking trio, with sends carrying the weight for non-follower reach. No published band exists — use your own median

    Two notes on those bands. They are directional priors from vendor analyses, not measurements from your account, so the number that actually governs your stop rule is your own format's history. And completion rate is length-dependent by construction, which is why a format with a drifting runtime cannot be tracked on it at all. Fix the length band first or the metric is noise.

    Plot the decay curve on episode number, not date

    Open a sheet. One row per episode. Four columns: episode number, publish date, the primary metric, and a rolling median.

    Plot the metric against episode number, not date. Cadence gaps are the most common reason a decay curve looks like a cliff when it is actually a two-week holiday. Episode number keeps the x-axis about the format.

    The shape you are looking for has three phases, and you need to be able to name which one you are in:

    1. Learning. First three to six episodes. Metric is below where it will settle, and noisy, because the audience does not recognise the shape yet. Nothing here is evidence of anything.
    2. Plateau. The metric stabilises in a band. This is the phase you are milking. It can run for a long time and it is boring to make. That is not a signal.
    3. Decay. A monotone downward drift across several consecutive windows. Not one bad episode. A drift.

    The distinction that matters operationally is between decay and variance. Short-form performance per post is dominated by feed variance: the same clip posted twice can land at wildly different reach.

    The rolling-median stop rule

    A median over the last three episodes (five if you post daily) throws out the single outlier that would otherwise trigger a false stop. A mean does not — one viral episode drags the mean up for three windows and hides a real decline.

    Write the rule down before you need it. A workable starting form:

    Retire the format when the 3-episode rolling median has sat below 80% of this format's own peak rolling median for three consecutive windows, and the direction across those three is down.

    Three things about that sentence are load-bearing:

    • This format's own peak, not your channel median. Formats have different natural levels. A 12-second This or That does not compare to a 90-second explainer, and forcing them onto one benchmark retires the wrong one.
    • Three consecutive windows. With a 3-episode window that is five or six episodes of evidence. It feels slow. It is calibrated against exactly the variance that makes single-pair conclusions worthless.
    • 80% is a placeholder. Set it from your own spread once you have twelve episodes: roughly one standard deviation below the plateau median is a reasonable anchor. The number matters less than having committed to one in advance.

    The rule also protects the opposite decision. If the rolling median is inside the band, the format is not spent, and the correct action when you are bored is to change the variable, not the format.

    Before you retire it, check three confounds

    A decaying curve is not proof the shape is dead. Rule these out first, in this order, because each is cheaper to fix than a new format.

    1. Slot drift. Did the publish time move? Early velocity varies by posting window, and a metric read at a fixed 24-hour checkpoint is not comparable across slots. Put it back and re-measure for three episodes.
    2. Length drift. Did episodes creep from 30 seconds to 70? On completion rate you have changed the denominator, not the performance. Re-band the length.
    3. Cold-open drift. Did the opening two seconds quietly change? The hook is the first frame, not the first line, and a format whose opening has drifted is no longer producing comparable episodes. Write the opening down and reuse it.

    Only if all three are clean is the decline about the shape.

    When it genuinely is, replace the variable before you replace the format. Most formats have one: the pairing, the question, the submission, the transformation. Swapping the variable is a fraction of the work of inventing a structure and it resets a good deal of the novelty. Eight repeatable short-video formats covers what those variables look like in practice. Retire the format only when a fresh variable has been through the learning phase and still sits under the band.

    Running the loop without adding work

    The tracking is the part that dies first. Three habits make it survive:

    • Log the metric at a fixed checkpoint — 24 hours, or 7 days, but the same one every time. Reading a metric at whatever hour you happen to open Studio is the fastest way to a curve that measures your schedule.
    • Log the episode number and the variable in the same row. Six months later you will want to know which variable was in the slot when it plateaued.
    • Keep the build identical between episodes. A saved edit reused as a template — the editor works off one re-renderable timeline, so saving an edit as a reusable draft keeps the caption style, safe zones and length band from drifting on their own. The 480p preview pass is free and carries a short per-user cooldown; a single charge applies to the final export only.

    The agent can check social performance on connected accounts. The sheet is still the thing that makes the decision.

    FAQ

    How many episodes before the curve means anything?

    Six, minimum, published on a fixed cadence. Below that you are inside the learning phase and measuring your own production quality rather than the format. Daily posting makes six a week; weekly posting makes it six weeks.

    What if one episode goes viral and wrecks the chart?

    That is what the median is for. A single outlier moves a 3-episode median far less than it moves a mean, which is the entire reason to use one. Log the outlier, keep it in the series, and do not treat it as the new peak — a peak set by one lottery win makes every subsequent episode look like decay.

    Can I run a decaying format and a new one at the same time?

    Yes, but only after the first is genuinely automatic, which usually takes six to eight episodes of practice. Running two unfamiliar formats reintroduces the weekly "what do we post" decision the format existed to remove. Different day, different slot, separate sheets.

    Does this apply to paid creative too?

    The stop logic does, the metrics do not. Paid creative decays against frequency and auction cost rather than completion, and the useful signals fire earlier. Creative fatigue and refresh cycles for ad accounts covers the paid version. The shared idea is the same: compare the asset to its own trailing peak, not to an account average.