How many episodes before you kill a series
Commit a pilot length, one metric and a written stop rule before episode one, so two noisy duds cannot cancel a series that has not been measured yet.
You cancelled the series after episode three because two and three "didn't hit." That is not a decision. It is a coin flip with extra meetings. Short-form reach is dominated by feed variance: the same file posted twice can land in completely different pools. The first handful of episodes are also the ones where viewers have not learned the shape yet. Killing inside that window is how a channel collects abandoned pilots and never owns a format long enough to have a median.
The fix is boring, and it has to happen before episode one: a committed pilot length, one metric, and a stop rule written down while you still like the idea. After that the series runs until the rule fires. Mood is not a clause.
Why two duds are not a verdict
Three things pile up in the early episodes, and none of them is "the format is dead."
The audience has not learned the shape. Recognition is part of why a series works at all. After a few episodes, regulars identify the cold open before the first line lands, which is a retention head start every later episode inherits. The video series recognition argument is exactly this: the unit of memory is the repeating shape, not the individual file. Episodes one to four are structurally not the peak. They are the tuition.
The feed lottery is louder than the format. A single-pair test on short-form is mostly noise. That is why a serious hook test batches variants and pre-commits a sample size instead of calling a winner on Tuesday night. A series cancelled on two weak posts is the same error with a sadder caption.
You are bored on a different clock than the audience. Your fatigue climbs with how many times you have built the thing. Their response climbs with how many times they have seen it. Those curves do not peak together. Episode five or six, when the caption is muscle memory, is the usual kill moment and is often still inside the audience's learning window.
If the series is interchangeable from episode to episode, that is a different problem: audience fatigue is a format problem, not a model one. A written stop rule still stops you confusing "this render is tired" with "this shape never worked."
The pilot you commit before episode one
A pilot is not "we'll see how it feels." It is a number of episodes you will publish at a locked slot, in a locked length band, before anyone is allowed to call it.
The length of that run is arithmetic, not inspiration.
- Budget a learning block. Viewers need several repetitions before first-frame recognition shows up. Treat the first four episodes as unscored for the kill decision. You still publish them. You do not let them vote.
- Budget a measurement block. A three-episode rolling median needs more than three rows. Two consecutive windows after the learning block is the minimum that is not just one outlier wearing a trench coat. That is six more episodes.
- Add them up. Four unscored plus six scored is ten. If you post daily, the learning block is shorter in calendar time, but you still want the same episode count. If you post weekly, ten episodes is ten weeks, which is exactly why people cheat and kill at week three.
Ten is a working default, not a law of nature. Write the number you will actually honour. An eight-episode pilot you finish beats a sixteen-episode one you abandon because "the vibe was off" at episode five.
Lock the production shape in the same sitting. Repeatable series formats exist so the per-episode job is filling a slot, not reinventing a structure. If the cold open, runtime band and caption preset are still being redesigned in week three, you have a sequence of one-offs with a shared title. The series catalogue is a place to steal a finished shape rather than debate one.
Cadence is not the kill lever. Do not raise posting rate to "give it a better chance." A weekly slot you can hold is a better pilot than a three-times-a-week slot that slips, because a slipped slot makes the 24-hour checkpoint incomparable. If the series is meant to run itself after the pilot, put the slot on a real schedule. Cron for a recurring series is "same day, same hour" as an actual job.
One metric, locked in writing
Choosing the metric after episode six means choosing the number that agrees with the decision you already made. Write it on the same page as the pilot length.
Pick the number that says whether people are still watching this shape, not whether one file got lucky with distribution. Views are the worst candidate.
| Surface | Metric to lock | How to read it |
|---|---|---|
| YouTube Shorts | Viewed vs. Swiped Away, Shorts Feed tab in Studio | Practitioner bands from Paddy Galloway's 2023 analysis of 3.3 billion views sit roughly 70–90% for a healthy clip, with collapse under 60%. Use your own median as the actual bar |
| TikTok | Average watch time, or completion rate | Length-banded floors circulate in vendor analyses: above about 50% average watch under 30 seconds, 40% from 30–60, 30% past a minute. They are not TikTok-published thresholds |
| Instagram Reels | Sends per reach | Adam Mosseri has named watch time, sends per reach and likes per reach as the ranking trio. Sends carry non-follower reach. There is no published healthy band, so use your own |
Two constraints, or the metric is junk.
Fix the length band. Completion rate changes denominator when runtime creeps from 28 seconds to 55. A series whose episodes wander cannot be killed or saved on completion, because you are not measuring the same object.
Read at a fixed checkpoint. Twenty-four hours is enough for a directional read on short-form and short enough that you will actually log it. Do not mix 24-hour numbers with seven-day numbers in the same column.
Open a sheet. One row per episode: number, publish timestamp, the metric, a three-episode rolling median. Plot against episode number, not date. A two-week holiday looks like a cliff on a date axis and like a gap on an episode axis. You care about the format, not the calendar.
The stop rule that survives a good week
A stop rule for a pilot is a different sentence from a retirement rule for a format that already had a plateau. You do not have a peak yet. You are asking whether the series ever earned the right to continue.
Write this before episode one, then do not edit it because episode seven popped:
After the committed N episodes, continue the series only if the three-episode rolling median of [metric], read at 24 hours, has sat at or above this channel's median for the same length band in at least two of the last three windows. Otherwise stop. Do not extend the pilot to "give the next hook a chance."
What that sentence is doing:
- Channel median for the same length band, not a screenshot from someone else's niche. A 12-second punchline format and a 50-second explainer do not share a bar.
- Two of the last three windows, not "the best episode." One viral file will drag a mean for a month. A median over three, required twice, ignores it.
- No extension clause. The whole point of a pilot is that the sample size is chosen while you are still optimistic. Extending it after a miss is how optional stopping sneaks back in.
If the rule says continue, you now have a series. The next question is when that run is spent: a decay question on a rolling median against this series' own peak, not against the channel. That is a later sheet. Do not smuggle it into the pilot.
If the rule says stop, stop. Do not rebrand the same skeleton with a new title and reset the counter. Log the kill: format, N, the median you got, the bar it missed. Then pick a different shape.
YouTube's Shows feature (seasons with custom artwork, grouped as bingeable units) is packaging for a series that survived a pilot, not a reason to keep one the numbers already failed. Numbered parts keep the backlog navigable. The claim that finishing Part 1 automatically queues Part 2 without search is widely repeated; treat returning viewers on episode N+1 as the thing you can measure.
FAQ
Can I kill it earlier if the comments are brutal?
Comments are a sample of people who stayed long enough to type. They are not a substitute for Viewed vs. Swiped Away or completion. If the metric is in band and comments hate a recurring bit, change the variable inside the format, not the format. If the metric is already dead, you still wait for the committed N unless the slot is operationally impossible: presenter gone, source material gone, account restricted. Those are production stops, not taste stops.
What if episode one is a runaway hit?
Then you have a lucky file and an untested series. Run the rest of the pilot. A rolling median exists to ignore outliers in both directions. Using a hit to skip remaining episodes is how people scale a shape they cannot reproduce.
Should the metric be subscribers, or revenue?
Not for the kill. Subscribers lag and mix in people who arrived for a different video. Revenue mixes in the offer and the landing page. The series question is whether people are watching this shape. If a series holds attention and does not convert, that is an offer problem sitting on a working format.
How does this change if I am running two series at once?
Give each series its own sheet, its own median, its own N. Do not average them. A dead series will hide inside a live one if they share a column. If you cannot staff two pilots at the locked slot and length, cut to one. The measurement loop has to be visible before you add volume.