How to Make Top-10 Videos With AI
How to make top-10 videos with AI: entry research, batch B-roll per item, countdown pacing that holds retention to #1, narration, and thumbnail strategy.
The top-10 format is the assembly line of faceless YouTube — the same chassis every episode, with only the cargo changing. That's exactly why it suits AI production: once you've built the machine (intro template, per-item structure, narration voice, B-roll pipeline), a new episode is research plus generation plus assembly, repeatable in a few hours. It's also why most top-10 channels fail: they build the chassis and forget that the ordering of the list is a retention instrument, not a formality.
Here's the full recipe — research standards, the per-item beat structure, batch B-roll generation, and the countdown pacing math that gets viewers to number one.
Research: 10 entries, 30 facts, zero filler
A top-10 video is only as good as entry number 7 — the middle of the list is where weak research shows. Set a research standard per episode: for each entry, collect three specific, verifiable facts (a number, a date, a comparison) and one "hook fact" surprising enough to open the segment. Thirty facts per episode sounds heavy; it's about 90 minutes of focused research, and it's the difference between a channel that gets cited and one that gets corrected in the comments.
Choose topics where visual variety is inherent — "10 abandoned megaprojects," "10 deepest places on Earth," "10 most expensive materials" — because each entry then demands visually distinct B-roll, which keeps the screen fresh for eight minutes. Topic selection is niche selection; the demand map in 20 faceless YouTube niches with AI demand is a good starting grid for lists that people actually search.
Script to a fixed per-item beat
Every entry follows the same internal shape, which is what makes the format bingeable — viewers learn the rhythm and settle in:
- Number card + name (2s): "Number 7: The Sarcophagus of Chernobyl."
- Hook fact (5s): the most surprising thing, first.
- Body (25–40s): the three researched facts, arranged as setup → detail → implication.
- Button (3s): one-line kicker that hands off to the next entry.
At 35–50 seconds per entry, ten entries plus a 20-second cold open and a 30-second finale lands the episode at 7–10 minutes — the top-10 sweet spot for long-form watch time. Write the whole script before generating anything; the script is the shot list.
Two scripting rules that move retention numbers: tease the #1 entry in the cold open ("and number one is worth more than the other nine combined"), and place your second-strongest entry at #6 or #5 — the mid-list slump is where viewers bail, so spend quality there, not just at the podium positions.
Generate B-roll in batches, one entry at a time
Each entry needs 4–8 shots to cover 35–50 seconds of narration without a single visual overstaying. That's 40–80 clips per episode — which sounds absurd until you batch it. Work entry by entry: write a shot list from the script's nouns ("aerial of the concrete arch," "close-up of rusted machinery," "1986 archival-style street scene"), then generate the batch with an AI B-roll generator, keeping a consistent visual grade across the episode by repeating the same style suffix in every prompt ("documentary tone, muted color, slight film grain").
Speed-optimized models earn their keep here — B-roll is volume work, and a speed-tier engine like Hailuo 2.3 Fast generates coverage shots quickly enough that a full episode's visuals are a single working session. Reserve slower, higher-fidelity generations for each entry's opening shot — the one that plays under the number card — and let fast models fill the connective footage. Mix in generated stills with slow push-ins for facts and figures; a 100% moving-footage episode is actually more fatiguing than one that breathes between motion and considered stills.
Narration and pacing: the countdown is a retention curve
Generate the voiceover with a consistent AI text-to-speech voice — authoritative, mid-energy, slightly faster than conversational, because list content tolerates pace. Generate per-entry rather than as one long read, so a script fix at #4 doesn't mean regenerating the episode.
Then pace the edit against the countdown's psychology:
| List position | Viewer state | Pacing move |
|---|---|---|
| 10–8 | Sampling, deciding to stay | Fastest cuts, strongest hook facts |
| 7–5 | The slump zone | Deploy your hidden-gem entry, vary shot types |
| 4–2 | Committed, anticipating | Slow slightly, build stakes per entry |
| 1 | Payoff | Longest entry, best visuals, callback to the cold open tease |
Add progress markers — the number card doubles as one — and word-timed captions for the key facts. YouTube chapters per entry help search and let skimmers jump, which counter-intuitively raises average view duration, because skippers who'd otherwise leave instead sample three entries.
Package: the thumbnail is half the video
Top-10 lives and dies on click-through. The winning thumbnail formula for lists: one dramatic image representing the #1 or most curiosity-inducing entry (not a collage of ten tiny images — those die at feed size), plus a short overlay like "#1 shocked me." Generate 3–4 candidates with an AI thumbnail generator and test alternates on underperforming episodes. Title with the keyword phrase people search ("Top 10 Deepest Places on Earth") — clever titles lose to searchable ones in this format, since lists are a search-driven genre. That also means episodes accumulate views for years, which is the format's quiet economics: a 50-episode back catalog becomes a compounding asset. The full channel-economics picture is in building a faceless YouTube business.
The weekly production loop
A sustainable one-person cadence:
- Day 1: Topic selection + research grid (30 facts).
- Day 2: Script to the per-item beat; shot lists fall out of the script.
- Day 3: Batch B-roll generation, entry by entry; generate narration.
- Day 4: Assembly — number cards, captions, music bed ducked under voice.
- Day 5: Thumbnail candidates, title, chapters, publish; cut two vertical excerpts (a single strong entry each) as Shorts that point to the full video.
One episode per week, each strengthening a searchable back catalog. The format's boring predictability is precisely its power — the machine improves every cycle while the audience learns to trust the rhythm.
FAQ
How long should a top-10 video be?
Seven to ten minutes: 35–50 seconds per entry, a 20-second cold open that teases #1, and a brief finale. Shorter than 6 minutes undercuts watch time; past 12, mid-list retention decays unless every entry is exceptional.
How many AI clips do I need per episode?
Plan 4–8 shots per entry — roughly 40–80 clips per episode. Batch-generate with a fast model using a consistent style suffix for visual coherence, spend higher-quality generations on each entry's opening shot, and mix in stills with slow push-ins to vary the texture.
Should the countdown go 10-to-1 or 1-to-10?
Ten-to-one, always — the countdown is the retention device, because the promised payoff sits at the end. Protect the mid-list slump (#7–5) with your second-best material, and make #1 the longest, best-produced entry so the payoff lands.
How do I pick top-10 topics that get views?
Pick searchable superlatives with inherent visual variety — deepest, most expensive, most remote, most dangerous. Check that people search the exact phrase, that ten genuinely distinct entries exist, and that each entry suggests obvious imagery; lists are a search genre, so the title keyword is the demand signal.
Can top-10 channels work fully faceless?
Yes — it's one of the most proven faceless formats: TTS narration, generated B-roll, number-card graphics, no presenter needed. The moat isn't the face; it's research quality and packaging consistency across a growing back catalog.
Build the chassis this week — one script template, one voice, one style suffix — and ship episode one. Versely's B-roll generation, TTS, and thumbnail tools run the whole line: start at the AI B-roll generator.