Score hooks at 24 hours, not at 7 days
Lifetime views are dominated by later distribution, not by your hook. A fixed 24-hour checkpoint sheet that reads average watch time instead of view counts.
You published two versions of the same video with different openings. A week later one has 41,000 views and the other has 9,000. You now know which post did better. You still do not know which hook is better, and if you promote the winner's opening into a production default you may well be promoting the loser.
That gap between "which post won" and "which hook won" is the whole problem with scoring creative on lifetime views. It is not a small effect and it does not average out over a handful of tests, because the thing that produces the variance is not random — it is a distribution mechanism that kicks in after the hook has finished doing its job.
What a view count is made of
Decompose a seven-day view count and the hook is one term among several:
- The initial seed. How many people the platform showed it to before it had any performance data. This varies with account state, posting time, and factors nobody outside the ranking team can enumerate.
- The opening's hold rate. The hook. The thing you were trying to measure.
- The body's hold rate. Everything after the promise lands.
- Sharing and off-platform pickup. Which is about transferability, not about the first two seconds.
- Later distribution waves. A post that catches a second wave on day three can multiply several-fold on a signal chain that started days after anyone saw your opening.
Only the second item is your variable. The last item is usually the largest, and it is the one that compounds — which means the longer you wait, the smaller a share of the final number your hook is responsible for. Waiting for more data makes the read worse, not better. This is the opposite of how sample size normally works, and it is why the instinct to "let it settle" is wrong for this specific measurement.
Why 24 hours
A checkpoint is not a claim that the video is finished. It is a decision to stop measuring while the thing you care about is still the dominant term.
Twenty-four hours is the conventional stopping point for three reasons. It covers a full daypart cycle, so a post published at 5pm and one published at 5pm the next day have both been exposed to the same rhythm of audience availability. It is long enough that the initial seed has been spent and the platform has made its first real judgement. And it is short enough that most second-wave distribution has not yet happened, which is exactly the term you are trying to exclude.
It is a convention rather than a derived optimum. If your account reliably gets its distribution decisions made in six hours, use six. The requirement is that the number is fixed in advance and identical across variants — a checkpoint you move is not a checkpoint, and comparing a 24-hour read against a 40-hour read is comparing two different measurements.
The metric, in priority order
Read the most direct opening-quality signal your platform exposes:
| Platform | Read this | Why |
|---|---|---|
| YouTube Shorts | Viewed vs. Swiped Away, in the Shorts Feed tab of Studio | The most direct opening signal any platform publishes. Practitioners treat roughly 70–90% as healthy and below 60% as trouble — a creator-side convention, not a YouTube threshold. |
| TikTok | Average watch time in seconds, then completion | Seconds do not move when clip length moves; completion needs a duration band to mean anything. |
| Reels | Average watch time | The first of the three signals Mosseri named. Read sends per reach separately; it is a transferability measure, not a hook measure. |
What you do not read is views. What you also do not read is likes, comments, or follows, all of which are downstream of the whole video and of your existing audience's habits.
The checkpoint sheet
One row per post, filled in twice: once at publish, once at the checkpoint. Nine columns, and every one of them is there because leaving it out has caused a bad conclusion at some point.
| Column | Filled at | Note |
|---|---|---|
| Post ID / link | Publish | |
| Pair ID | Publish | Which test this belongs to. Blank means it is not a test. |
| Variant | Publish | A or B. Decide before you publish which is which. |
| Hook family | Publish | Loss, curiosity, visual demo, direct callout, contrarian. The thing you will aggregate on later. |
| Clip length | Publish | Because completion is meaningless without it. |
| Published at | Publish | Local clock time, not just the date. |
| Checkpoint due | Publish | Published-at plus 24 hours. Write it down so you cannot drift. |
| Primary metric | Checkpoint | From the table above. One number. |
| Secondary metric | Checkpoint | Completion, or sends per reach. Recorded, not decided on. |
The two disciplines that make the sheet work are both about writing things down before you can be influenced by results. Pick the metric before you look — deciding after you have seen the numbers means picking whichever metric favours the answer you already liked. And write the checkpoint time at publish, because the failure mode is not forgetting to check, it is checking early, seeing a result you like, and treating that as the reading.
Peeking is the dominant failure of creator-run tests, and it has a specific signature: early data skews toward whichever variant posted first, and reverses with volume. If you look at four hours you will usually see variant A ahead. That is a scheduling artefact.
Three rules around the sheet
Never post variants simultaneously. Two versions of the same video live at once compete for the same audience pool and cannibalise each other's seeding. The standard is the same clock time on consecutive days.
Run at least three pairs before concluding. A single short-form pair is dominated by feed-lottery variance. Three pairs pointing the same direction is a finding; one pair is a story. This is the same requirement behind any creative test worth running, and it is why the sheet has a pair ID column — the unit of analysis is the pair, not the post.
Batch the production so the cadence is bearable. Three pairs is six posts. If each one is a bespoke build you will run one test and never run another. Generating variants in a batch and reusing the body is what makes the protocol survive contact with a real schedule; the argument for testing twenty creatives at once applies in miniature here. Scheduling the pairs in advance also removes the temptation to publish both on the same day when the week gets busy — the agent can post or schedule to social media across connected accounts, which fixes the timing without you having to be at a desk at 5pm two days running.
Where 24 hours is the wrong window
Two cases, both worth knowing so you do not apply this rule where it fails.
YouTube long-form packaging. YouTube's native title and thumbnail test runs up to three variants and picks its winner on watch time rather than click-through, resolving within about two weeks — faster at high impression volume. Do not read it early. Shorts, Made for Kids, private, mature-audience, and age-restricted videos are ineligible (a test on an age-restricted video will not run), as are scheduled lives and Premieres until they convert. Changing the title or thumbnail outside the test stops it. Different tool, different window, different metric.
Anything you are measuring for conversion rather than attention. A hook checkpoint at 24 hours tells you about attention. If the question is whether a creative sells, the answer arrives on a purchase cycle, not a distribution cycle, and 24 hours is far too early.
Everything else — hooks, openings, first frames, opening line rewrites — reads better at a fixed early checkpoint than at lifetime, and the accounts that switch usually find their previous conclusions were about seeding rather than about creative. That is uncomfortable and also the point: the sheet exists to stop you learning things that are not true. It pairs naturally with the hook testing framework, which handles what to test; this handles when to stop looking.
FAQ
What if a post is still climbing at 24 hours?
Record the number anyway and let the post keep running. You are not deciding whether to keep it; you are logging a comparable measurement. A variant that is still climbing at the checkpoint is interesting information about distribution, and it belongs in the notes column rather than in a decision to extend the window for that one post.
Does the checkpoint have to be exactly 24 hours?
It has to be the same for both variants and fixed before publishing. Twenty-four is convenient because it controls for time of day automatically. If you check at 26 hours because you were asleep, check the other variant at 26 hours too and note the drift.
Can I use views if I have no access to watch-time data?
If watch time genuinely is not available, hook rate or three-second view counts are the next best proxies, since they are still measured close to the opening. Total views should be the last resort, and if you use them, tighten the checkpoint rather than loosening it — a six-hour view count is much more about the hook than a seven-day one is.
How long before the sheet is worth anything?
About six weeks at a normal cadence, which is roughly six pairs plus the non-test posts filling in the hook family column. At that point you can aggregate by hook family across everything you published rather than only across formal pairs, and that aggregate is usually more useful than any individual test result, because it covers your real content mix rather than the narrow slice you chose to test.