A creative test log you'll actually read
Tests without a written record get re-run every quarter at full cost. Log these fields per test, then review on a rhythm that feeds a hypothesis backlog.
Tests without a written record get re-run and re-argued every quarter at full cost. The cost is not only the generation. It is the distribution slots, the weekday you burned, and the meeting where two people remember opposite winners. A log that nobody opens is decoration. A log with twelve fields that take forty minutes to fill is decoration with extra steps.
The version that gets read is a row per test, filled on the day you ship and closed at the checkpoint you already named. The review rhythm is what turns those rows into a hypothesis backlog instead of a graveyard of screenshots.
Fields worth logging, and nothing else
If a field does not change a future decision, drop it. This is the set that has survived contact with real sheets.
| Field | What to write | Why it exists |
|---|---|---|
| Test ID | 2026-08-20-tt-hook-problem-vs-curiosity |
Searchable; date plus surface plus variable |
| Platform / account / format | TikTok, main, 9:16, ~30s | Results do not pool across these |
| Hypothesis | One reusable sentence | If you cannot write this, you do not have a test |
| Variable | The one row that was allowed to change | Catches confounded pairs after the fact |
| Locked | Caption, sound, hashtags, cover, clock, body | The spec, compressed |
| A / B | Five words each | So a stranger can read the row |
| Metric + checkpoint | "24h AWT" or "VVSA at 24h" | Named before either post goes up |
| Stop rule | "Ship the lead if obvious with names hidden; else tie" | Stops peeking from rewriting history |
| Posted | A Tue 17:00 / B Wed 17:00 | Weekday confound lives here |
| Result | Numbers at the checkpoint, same age | Not views, not a screenshot without transcription |
| Decision | Ship family / kill / tie / spoiled | One word you will obey |
| Next hypothesis | The sentence this result produced | This is the backlog |
A filled row looks like this, not like a paragraph:
ID: 2026-08-12-shorts-hook-demo-vs-talking-head. Hypothesis: Visual-demo openings beat talking-head claims on 24h viewed-vs-swiped for the tutorial series. Locked: same body, caption, sound, hashtags, cover. A: hands-only demo of the method. B: talking head, same line. Posted: A Wed 18:00, B Thu 18:00. Result: A leads B on VVSA at 24h by a gap you would still pick with the names hidden. Decision: ship A as the family rule, do not ship this wording. Next: same family on two more topics before we call it.
If the "next" cell is blank, you logged a souvenir.
What you do not log:
- Raw views as the result, unless views were the named metric (they should not have been, for hooks).
- A paragraph of feelings. The sheet gets numbers and a decision.
- Every export setting. Lock them, unless a spoiled test was caused by one (wrong crop, missing captions). Then write "spoiled: no burn-in on B" and rerun.
- Screenshots as the only record. Transcribe the three numbers you named.
If you already track a wider creative analytics set, the test log is not a second dashboard. It is the decision layer on top. Dashboards describe what happened to posts. The log describes what you were trying to learn.
A review rhythm that produces the next test
A log that is only written when you remember is not a system. Put two slots on the calendar. Protect them the way you would protect a publish time.
Weekly, 20 minutes, same weekday. Open every test whose checkpoint has arrived. For each one: transcribe the result, write the decision, write the next hypothesis, close the row. Do not redesign the channel in this slot. Do not start a new argument about a test that has not hit its stop. If you pull account performance here, you pull it for the rows that are due, not for every live post.
Three outputs, and then you stop:
- Closed rows (result + decision).
- At most one new pair scheduled, copied from a "next hypothesis" cell.
- A list of spoiled tests to rerun, if any.
If you leave with seven "we should" notes and zero closed rows, you had a conversation. The slot failed.
Monthly, 30 minutes. Read the closed rows as a set, not as anecdotes. Count win rates by family, by platform, by weekday. This is where a Tuesday bias shows up, where "visual demo" is 4–1 on Shorts and 1–3 on Reels, and where you notice you have run the same caption test twice because nobody searched the sheet. Audience retention reviews still own the curves. This monthly pass owns the questions those curves keep raising.
The monthly pass has one job: turn repeating results into rules, and repeating arguments into tests that are actually on the calendar.
- A family that won three pairs becomes a default for the next month's production, not a memory.
- A variable that tied twice gets dropped. Tie means "not where the leverage is." Stop spending slots on it.
- A result you cannot find is a result you will pay for again. Search before you greenlight a "new" idea.
This is also where creative fatigue belongs as a note, not as a vibe. If a family that was 4–1 last month is now losing to the same control, write "fatigue?" and run one more pair before you throw the family out. Do not skip the pair because someone is bored of the format.
The log is the hypothesis backlog
Most teams keep a backlog of ideas ("try a darker hook," "test a question caption"). Ideas are cheap and they do not remember what was already spent. A hypothesis backlog is a queue of sentences the log produced, each one tied to a previous row.
The conversion is mechanical:
| Closed decision | Backlog item it writes |
|---|---|
| Family A wins 2 of 3 | Two more pairs in family A on new topics, same control |
| Tie on caption first line | Kill caption tests this month; spend slots on hook family |
| A wins only on Tuesdays | Reverse-order rerun; do not ship the hook |
| Spoiled (cover mismatch) | Rerun the same pair, same weekdays, same files |
| Paid winner, never run organic | One organic pair of that family, consecutive days, not a cross-post dump |
Nothing enters the backlog unless a closed row pointed at it. Nothing leaves except by being scheduled as a pair or being killed in the monthly pass. "We should try negative hooks" without a parent row is how the log gets bypassed and the quarter repeats.
Production follows the backlog, not the other way around. If the next hypothesis is "same family, new topic," you need another body and two heads, not a new brand film. Batch generation is for filling those rows. It is not a substitute for them. A 30-day batching calendar that ignores the log will happily produce thirty files of a family you already tied. Testing velocity is closed tests per week, which the log can count, not files exported.
Keep the sheet where the people who post can edit it. A Notion database, a Google sheet, a table in the same doc as the content calendar. Not a slide deck. Not a folder of screen recordings named IMG_8841. The owner of a row is the person who scheduled the pair. They fill "posted" the day it ships and "result" the day it stops. If that is always one person, the log dies when they are off. Rotate the close-out, not the format.
Search the ID column before you write a new hypothesis. If last May already killed "question versus claim in the caption" on this account, you do not get to spend August on it unless the monthly pass reopened it for a stated reason (new format, new audience, fatigue on the old winner). Re-runs without a reason are how full cost comes back.
FAQ
Spreadsheet or a document?
A table. Rows, not prose. A document turns into a narrative nobody skims. Twelve columns in a sheet, one row per test, frozen header. If you need a paragraph, you are writing a post-mortem for a spoiled test, and that paragraph lives in a notes column, not in Slack alone.
How far back do we keep rows?
Keep them as long as the account is the same account. A year of closed tests is how you stop repeating a tie. Archive, do not delete, when you change format or audience so thoroughly that old rows would mislead. Even then, keep the archive searchable. The cost of storing rows is nothing next to the cost of rerunning them.
Do paid and organic share a log?
Same workbook, different tab or a "surface" column that you never filter out by accident. A paid winner is a hypothesis for organic, not a pasted result. Mixing the numbers in one result cell is how a split test launders a lottery. The creative-testing playbook still applies to how each row is designed. The log is how those rows accumulate.
What if nobody fills the result column?
Then you do not have a log, and you should stop starting tests until the close-out is easier than skipping it. Make the weekly slot the close-out, not an optional review. A pair with no result at checkpoint is marked "unscored" and does not count as a win for anyone's favourite. Two unscored weeks in a row means you are posting, not testing. Cut volume until the sheet is current. Unscored tests are the expensive ones, because you paid for them twice: once to run, once to argue about without numbers.