How to Make a Mini-Documentary With AI
Make a mini-documentary with AI: interview-spine structure, an avatar narrator, cinematic b-roll cutaways, lower thirds, and a three-act edit under ten minutes.
A mini-documentary — 6 to 10 minutes, one subject, one arc — is the most credibility-dense format a brand or creator can publish. It's also traditionally the most expensive per minute: interview shoot, b-roll days, an editor finding the story in twelve hours of rushes. The AI recipe inverts that process: you write the story first, deliver the interview through an on-camera presenter you generate, and shoot your "b-roll days" as cinematic prompt sessions. This guide covers how to make a mini-documentary with AI using the interview-spine structure — the editorial skeleton nearly every modern doc uses, where a talking head carries the narrative and cutaways carry the atmosphere.
The interview spine: write the transcript before the film exists
Watch any streaming-era documentary with the sound off and you'll see the same skeleton: a person in a chair talking, interrupted every 10–20 seconds by cutaway footage. The talking head is the spine; the cutaways are the ribs. Traditional docs discover this spine in the edit by cutting down hours of interview. You're going to write it directly — which is faster and produces a tighter film, at the cost of the serendipity a real interview brings. That's the honest trade.
Write your spine as a spoken-word transcript, not an essay: 900–1,300 words for an 8-minute film, in the first person, with the specific texture real interviews have — hesitations rendered as sentence fragments, concrete details over abstractions ("The first loaf I sold was burnt on one side. She bought it anyway."). Structure it in three acts:
- Act I — The hook and the setup (90s): open mid-story on the most arresting line, then establish who/what/where.
- Act II — The struggle (4–5 min): the obstacle, the doubt, the turn. This is 60% of the film.
- Act III — The resolution and the meaning (90s–2 min): where things landed, and the one-sentence idea the viewer keeps.
If your subject is a real person — a founder, a craftsperson, a customer — interview them for real (even a voice memo) and cut that into the spine. Real voices beat written ones whenever you can get them; this recipe's generated-presenter path is for when you can't, or when the "subject" is a topic rather than a person.
The on-camera presenter: an avatar that holds the frame
The spine needs a face. A digital-twin avatar — built from footage of you or a presenter — delivers your transcript on camera with natural gesture and lip sync, which means retakes are free and pickups are a paste-edit. HeyGen Avatar V5 digital twin is built for exactly this seated-interview register.
Three direction choices make the avatar read as documentary rather than corporate explainer:
- Frame it like an interview. Subject offset to one third, looking slightly off-camera toward an unseen interviewer — not straight down the lens. Straight-to-lens is presenter grammar; off-axis is documentary grammar.
- Break the transcript into answer-length chunks. Generate 20–40 second segments, not one monologue — docs cut between "answers," and the micro-resets between segments mimic that texture.
- Keep the background honest and underlit. A workshop, a kitchen, a bookshelf in soft light. If the film is about a topic rather than a person, you can skip the avatar entirely and run voiceover-only — but the on-camera spine measurably outperforms pure narration for trust, which is the format's whole point.
Disclose the synthetic presenter in your description. Documentary is a trust genre; spending that trust on an undisclosed avatar is a bad trade every time.
Cutaways: shoot your b-roll days in an afternoon
Now the ribs. Every 10–20 seconds of spine gets interrupted by 3–6 seconds of cutaway, which for an 8-minute film means 25–35 shots — your biggest generation batch. Write the cutaway list against the transcript: go line by line and note what the viewer should see while each line plays. The craft rule is illustrate the noun or the feeling, never the whole sentence — for "the first loaf I sold was burnt," show hands scoring dough or an oven's glow, not a literal burnt loaf with a customer.
Generate cutaways in a consistent cinematic register with an AI b-roll generator: pick one grade ("overcast natural light, muted color, shallow depth, handheld drift") and repeat it in every prompt so 30 shots feel like one shoot. Mix three shot sizes deliberately — wide establishing, medium action, macro detail — because all-macro or all-wide is the fastest tell of assembled stock. And for the two or three hero cutaways at your act turns, spend top-tier cinematic model credits; the mid-film money shot earns its cost, while transitional shots can come from fast tiers.
| Cutaway type | Count for 8 min | Model tier | Duration |
|---|---|---|---|
| Establishing wides | 5–6 | Cinematic | 4–6s |
| Action mediums | 12–15 | Fast/standard | 3–4s |
| Macro details | 8–10 | Fast/standard | 2–3s |
| Hero shots (act turns) | 2–3 | Top cinematic | 5–8s |
The documentary edit: rhythm, lower thirds, and silence
Assembly is where doc grammar lives. Lay the spine down first — the full avatar interview in sequence — then cut cutaways over it, keeping the interview audio running underneath (that continuous voice under changing pictures is the single strongest "this is a documentary" signal). Let the presenter be on screen roughly 40% of the time: too much face is a monologue, too little is a slideshow.
Finish with the genre's furniture: a lower third the first time the presenter appears (name + role, small type, 4 seconds), a title card 20–30 seconds in — after the hook, never before — and section breathers at act turns: 3 seconds of a hero shot with music and no voice. Score it sparse: one ambient theme that drops out at the emotional peak of Act II and returns changed in Act III. Silence at the peak is the oldest documentary trick and it still works. Timed captions matter here too — docs get watched on lunch breaks with the sound off.
The same spine-and-ribs recipe scales down to a 60-second vertical cut: hook line, three cutaways, resolution line. Post it as the trailer that funnels to the full film. For the adjacent recipes — archival subjects need a different visual system entirely — see how to make history videos with AI, and for using multi-scene story arcs in brand films, brand storytelling with multi-scene AI movies.
FAQ
How do you make a mini-documentary with AI?
Write the interview transcript first — 900–1,300 words in three acts — then generate an on-camera presenter (a digital-twin avatar framed off-axis like an interviewee) delivering it in answer-length segments. Generate 25–35 cinematic cutaways keyed to specific lines, edit them over the continuous interview audio, and finish with lower thirds, a delayed title card, and a sparse score.
How long should a mini-documentary be?
6–10 minutes is the format's natural range — long enough for a three-act arc, short enough to hold social-referred viewers. Cut a 60-second vertical trailer from the same material using the hook line and your three best cutaways.
Can I use a real interview instead of an avatar?
Yes, and you should when you can — even a phone-recorded voice memo from a real subject beats written narration for authenticity, and the AI recipe still supplies every cutaway around it. The generated-presenter path exists for topic-driven films or when no subject can be on camera.
What makes AI b-roll feel like documentary footage?
One consistent grade repeated across every prompt, a deliberate mix of wide, medium, and macro shot sizes, and cutaways that illustrate a noun or feeling from the narration rather than literally restaging the sentence. Continuous interview audio underneath changing shots does the rest.
Do I need to disclose the AI presenter?
Yes — documentary trades entirely on trust, so disclose the synthetic presenter and any generated footage in the description at minimum. Platforms increasingly require AI disclosure anyway, and in this genre transparency costs you far less than discovery does.
Write your three-act transcript this week, then build the spine and every cutaway in one place with Versely's AI video generator — your first mini-doc can ship this month, not this quarter.