Place CTAs at 55-75% of runtime
End-screen CTAs reach the smallest remaining audience. Place the ask at 55-75% of runtime, right after a payoff, then A/B the timestamp.
An end-screen CTA is an ask served to whoever is still watching. That is the smallest remaining cohort in the video. The conversion gap between a mid-roll ask delivered right after a payoff and the same ask on the end card is usually not a copy problem. It is a placement problem.
The working window is 55–75% of runtime. The adjacency rule is stricter than the window: the ask sits immediately after a payoff, framed around the value you just delivered, not before the first proof and not as a leftover on the last frame.
Why the last card is the worst audience
Retention is a shrinking set. A retention curve that looks healthy in the middle still leaves you talking to a minority of starters by the last seconds. An end CTA is shown to the people who already stayed, which is the set you have the least of.
Two curve shapes make this worse than a gentle fade.
A mid-video cliff between about 40% and 70% of runtime is the signature of a tangent, an ad read, or a format break. A CTA that interrupts a payoff is that cliff. You will read it as "people don't like being sold to." The graph is usually saying "you asked in the middle of the sentence."
A padding cliff around the 5–6 minute mark on a stretched idea is the other trap. Teams pad a six-minute video to ten so there is room for a mid-roll. Viewers leave at the stretch, and the mid-roll you invented the runtime for never reaches them. Publish at true length and place the ask inside the video you actually have.
None of this means the last two seconds are wasted. They still have a job: a readable frozen frame, a loop that does not look like a glitch, and a line that matches the destination. That is end-frame design, not the only ask in the video. Use the exit as reinforcement. Do not use it as the first time you say what to do.
The 55–75% window and the adjacency rule
The window that keeps showing up in practitioner write-ups is 55–75% of runtime, immediately after a payoff. Treat any specific lift percentage as something to reproduce on your own channel, not a number to write into a forecast. Hosted-player studies and social-feed graphs do not always agree on whether the last card converts remaining viewers better; they do agree that the remaining audience is smaller. Wistia's 2026 State of Video, drawn from more than 13 million hosted videos, found interactive elements at the end of the file had the highest click-through among whoever was still watching. That is a hosted-player result, not a YouTube or TikTok ranking result, and it measures conversion of the people who remained, not how many people still are. Both facts can be true at once, which is why you A/B the timestamp instead of copying a vendor window into a forecast.
After a payoff means the viewer has just received the thing the title promised. In a tutorial, that is the moment the method works, not the moment you announce there will be a method. In a product demo, that is the after-state, not the unboxing. In a story, that is the turn, not the setup. The ask then frames the value just delivered: "If you want the checklist I just used, it is here," not "Smash subscribe" dropped into a random trough.
Never before the first payoff. Pre-value asks are a common cause of a step-down around 0:20, when the viewer has confirmed the packaging and has not yet been paid. You taught them the video is an ad before it was a video. Move the ask later. Do not make the ask quieter.
Worked timestamps, using the midpoint of the window as a starting mark and then sliding to the nearest payoff:
| Runtime | 55% mark | 75% mark | Put the ask after |
|---|---|---|---|
| 0:15 short | 0:08 | 0:11 | The one demo beat, if it exists; otherwise skip a mid-roll and design the exit |
| 0:30 short | 0:16 | 0:22 | The demo or the proof line, then a short ask, then a designed last two seconds |
| 0:60 | 0:33 | 0:45 | The result, not the recap |
| 8:00 long-form | 4:24 | 6:00 | The worked example that proves the promise |
| 10:00 long-form | 5:30 | 7:30 | The method working, before the recap and the subscribe sting |
A 12-second clip does not have a mid-roll in any useful sense. The whole video is the hook. Put a readable instruction on the freeze-frame and stop pretending you have a placement test.
The ask itself still follows the modular hook, body, CTA split: one action, spoken and shown, in language that matches the destination. Placement is when that module fires, not a new script.
How to A/B position without confounding the rest
You are testing the timestamp of the ask, not the offer, the overlay, or the voice line. If those also move, you will not know whether 6:10 beat 9:40 or whether the new wording did.
Hold this constant across the pair:
- The same spoken CTA line and the same on-screen text.
- The same destination.
- The same body cut, music, and captions up to the ask.
- The same end-frame, if you keep one as reinforcement.
Move only this:
- When the spoken line and the matching overlay appear.
Two positions are enough. After-payoff inside 55–75% versus the same ask on the end card. A third position (after the first proof versus after the second, both inside the window) is a later test, not a simultaneous one.
On an EDL timeline this is a slide, not a reshoot. Duplicate the sequence, drag the overlay and the voice clip to the new in-point, and check the join on the editor's 480p preview pass. That preview is free and carries a short per-user cooldown; a single charge applies to the final export only. If the join is ugly, you placed the ask on a cut rather than after a beat. Slide it to the hold, not to the splice.
Metric, named before you look: clicks, follows, or comments-of-intent per unique viewer who reached the ask, not raw views and not completion rate alone. Completion can fall a little when a working ask sends people out. That is sometimes the point. What you cannot tolerate is a cliff at the ask from viewers who have not yet been paid. Read the curve at the timestamp, not just the average.
For the overlay, keep text out of the bottom ~15% and the top ~10% of a vertical frame so platform UI does not sit on the instruction. Time it with a text overlay on the same beat as the spoken line. Sound-off viewing is common enough that a spoken-only ask is a placement test you already lost.
What changes on short-form versus long-form
Long-form has room. You can pay the viewer, ask, and still have two minutes of application left. YouTube's published recommendation goals include long-term viewer satisfaction, not only raw watch time. A mid-roll ask after the method can be structurally correct even if average percentage viewed dips: a viewer who got the answer and left is not the same signal as a viewer who bounced mid-sentence. You are not trying to trap people through a recap they did not need.
Short-form has almost no room, and ranking still cares whether people finish and send. On a 30-second clip, 55–75% is seconds 16–22. That is still mid-roll relative to the runtime. It is not "late." The pattern that works:
- Hook on frame one.
- Payoff as fast as the idea allows.
- One-line ask on the hold after that payoff.
- A last two seconds that can loop without looking broken.
Do not delete the designed exit because the mid-roll "won." The mid-roll is for the people who are ready now. The exit is for the people who will watch it twice and tap on the second pass. They are different jobs on the same timeline.
If you already read curves weekly, look at the 40–70% band specifically for a step that lines up with your current ask. Audience retention analysis is the shape list; this post is the placement rule that shape is usually pointing at.
FAQ
Will a mid-roll CTA destroy retention?
It will if you place it before the first payoff or on top of a format break. A short ask immediately after the method works usually shows up as a small step, not a cliff, because the viewer has something in hand. If the curve falls off a wall at the ask, you asked too early, too long, or in a different tone from the rest of the video. Shorten the line before you move it back to the end card.
What if the video is under 20 seconds?
You do not have a 55–75% test worth running. The runtime is the hook. Put the instruction on a still that survives a pause, keep it out of the UI safe zones, and spend the test slots on the first frame instead.
Should I drop the end-screen entirely once mid-roll wins?
No. Drop relying on it as the only ask. Keep a short reinforcement on the exit so the loop and the freeze-frame still tell people where to go. The test is "ask after the payoff, then remind" versus "ask only at the end," not "delete the last card."
How many positions do I test at once?
Two. After-payoff versus end card, identical copy. If you ship three timestamps plus a new overlay colour, you are not running a placement test. You are launching three videos and picking a favourite.