Hook, Body, CTA: The Ad Script Formula That Scales
The hook-body-CTA ad script formula broken down by the second: hook archetypes, proof structures, CTA patterns, and timing splits for 15, 30, and 60s ads.
Every video ad that has ever printed money for me fits on an index card: three seconds that stop the scroll, fifteen to forty seconds that make one argument, and five seconds that tell the viewer exactly what to do. Hook, body, CTA. It sounds too simple to be a competitive advantage, and that's exactly why it is one — most ad scripts I audit fail not because the idea is bad but because the structure is smeared. The hook makes two promises, the body makes four arguments, and the CTA arrives apologetically at second 52 after the viewer left at second 9.
The formula scales for a second reason that matters more in 2026: it's modular. When your script has clean joints, you can swap any section without reshooting the others — which is precisely what AI variant generation is good at. A structured script isn't just better copy; it's better raw material for testing.
The timing splits that actually work
Before archetypes, get the proportions right. The most common failure is a body-heavy script with a starved hook and a rushed CTA.
| Ad length | Hook | Body | CTA | Best for |
|---|---|---|---|---|
| 15s | 0–3s | 3–11s | 11–15s | Retargeting, price/offer ads |
| 30s | 0–3s | 3–24s | 24–30s | Cold prospecting, single-benefit stories |
| 60s | 0–5s | 5–50s | 50–60s | Considered purchases, founder stories, VSL-lite |
Notice what doesn't change: the hook is always ~3 seconds (5 at most), regardless of total length. Length is added to the body, never to the opening. If your product needs 60 seconds of persuasion, it still gets 3 seconds to earn them. Whether it should get 60 seconds at all is a data question — I dug into the retention numbers in Video Ad Length: What the Data Says in 2026.
Hooks: six archetypes, not infinite creativity
Hooks feel like the creative part, but in practice almost every performer falls into a small set of archetypes:
- Callout. Name the viewer: "If you run Facebook ads for a Shopify store..." Ruthlessly effective for niche products because it buys relevance instantly.
- Contrarian claim. "Your moisturizer is why your skin is dry." Earns the next ten seconds by creating a debt: now you have to explain.
- Result-first. Show the after, then rewind. "This took four minutes" over a finished result.
- Pattern interrupt. A visually wrong or unexpected first frame — product in an absurd context, an unexpected motion. Highest ceiling, highest miss rate.
- Question with a cost. "How much did you pay for your last video ad?" Works when the honest answer hurts.
- In-medias-res. Start mid-story: "So the second batch sold out too, and I need to explain why." Borrowed from creator content; reads native.
Two rules across all six. First, the visual and the verbal hook must be different information — if the voiceover says "this cleared my skin," the frame should show the skin, not a person saying the sentence. Second, write hooks last but shoot them first-class: the hook deserves your best generation, your best take, your best frame. If you want raw material, there's a library of 50 tested openers in the AI hooks post.
The body: one argument, three proofs
The body's job is to pay off the hook's debt with a single argument. Not three benefits — one benefit, supported three ways. The structure I use for nearly everything:
- Mechanism (why it works). One sentence on what makes the product different. "The strap anchors under the shoulder blade, not the neck." Viewers don't need the science; they need to believe there is a science.
- Demonstration (show it working). The product doing the thing, in real or generated footage. For physical products this is the section most worth spending generation credits on: a clean demo shot beats any adjective.
- Social or personal proof. One line: a number ("12,000 five-star reviews"), a real quoted review, or the presenter's own before/after. Keep it to one; stacking proof reads as insecurity.
Then stop. The most disciplined thing a script can do is not add the fourth point. If you have a second argument, that's a second ad — and with AI generation, the cost of making it a second ad is nearly zero, which removes the last excuse for cramming.
For long-form persuasion (60s+), the same skeleton stretches into problem-agitate-mechanism-proof; that's effectively a compressed VSL, and the AI VSL guide covers the extended version.
CTAs: specific beats clever
CTA weakness is the most fixable problem in ad scripts. The patterns that consistently convert:
- Action + destination + reason: "Tap Shop Now to get the starter kit — the 40% launch price ends Friday." Three pieces of information, five seconds.
- Risk-reversal close: "Try it for 30 days; if you don't sleep better, we refund you." Best for considered purchases.
- Continuity close: "Follow for part two" — only for content-led accounts, never for direct response.
What kills CTAs: vagueness ("check us out"), multiple asks ("visit the site, follow us, and use code..."), and burying the ask in the last half-second. Give the CTA its full time slot and an on-screen caption that repeats it — a large share of viewers are watching muted, and a spoken-only CTA doesn't exist for them.
Why this formula is built for AI production
Here's where structure becomes leverage. A hook-body-CTA script is three independent modules, and modern AI video tooling treats them that way:
- Hook swaps. Generate four different 3-second openers against the same body — different archetype, same promise. This is the single highest-ROI test in most ad accounts, and with an AI video generator each new hook is minutes, not a shoot day.
- Body consistency. Reference-to-video models keep the same presenter and product across every module, so a swapped hook doesn't look stitched on.
- CTA localization. Same ad, five offer endings — one per market or promo window — without touching the argument.
- Full assembly. For talking-head formats, the UGC video generator takes a script broken into these three parts and returns the presenter, product overlay, and timed captions in one pass.
Write once, structured; generate many. That's the whole trick.
A worked example: 30-second UGC ad
To make it concrete, here's the skeleton filled in for a fictional posture trainer:
Hook (0–3s): [Presenter mid-wince, hand on neck] "My desk job was destroying my neck — until I fixed the thing nobody talks about."
Body (3–24s): "It's not your chair. It's that your shoulders roll forward the moment you stop thinking about them. [Mechanism] This trainer vibrates the second you slouch — it doesn't hold you up, it reminds you, so your own muscles do the work. [Demo: product on, slouch, buzz, correction] I've worn it two weeks; [Proof] my afternoon headaches are just... gone."
CTA (24–30s): "Tap Shop Now for 30% off your first one — and if your posture isn't visibly different in 30 days, they refund you."
Count the moves: one callout-adjacent hook, one mechanism, one demo, one proof, one risk-reversed CTA. Nothing else. Every sentence is load-bearing, and every module can be swapped independently when the ad fatigues.
FAQ
What is the hook-body-CTA formula?
It's the three-part structure behind most direct-response video ads: a 3-second hook that stops the scroll and makes one promise, a body that pays off that promise with one argument and up to three proofs, and a 5-second CTA that gives a specific action, destination, and reason to act now.
How long should the hook be in a video ad?
Three seconds, five at the absolute most — regardless of total ad length. Extra runtime should always go to the body, never the opening. Most platforms report drop-off steepest in the first three seconds, so the hook's only job is to buy the next ten.
How many benefits should a video ad script cover?
One. A single argument supported by a mechanism, a demonstration, and one piece of proof outperforms multi-benefit scripts in nearly every test I've run. If you have a second strong benefit, make a second ad — AI generation makes the marginal cost of that trivial.
Can AI write a good ad script?
AI drafts good structured scripts when you constrain it to the formula: ask for a specific hook archetype, one argument, and a specific CTA pattern, then edit for voice and claims. Where AI clearly wins is variation — producing eight hook options for a body you've already validated.
Do CTAs really need to be spoken and on screen?
Yes. A meaningful share of feed viewers watch muted, so a spoken-only CTA is invisible to them, and a caption-only CTA is weaker for sound-on viewers. Do both, and give the CTA its full time slot instead of squeezing it into the final half-second.
When your next script is written, feed it straight into the AI video generator — hooks, bodies, and CTAs as separate generations you can remix all quarter. Free credits daily.