Inside the digestive system explainer workflow
Eleven scenes, one character reference used twice, and 396 credits. A scene-by-scene teardown of a published explainer recipe, and how to fork it.
The Digestive System Explainer is eleven scenes, four seconds each, 44 seconds of finished 9:16 video, animated with Kling 3 Turbo Image to Video, and listed at 396 credits at preview resolution or 495 at full. A kid-friendly journey: a boy bites a cheeseburger, the camera follows the food all the way through, and the boy comes back at the end to give a thumbs up.
It is worth taking apart not for the subject but for the structure, which is portable. Strip out the anatomy and what is left is a reusable shape for explaining any process with a beginning, a sequence of stages, and an end, which is most explainers anyone actually needs to make.
The shape: a bookend around a journey
Read the eleven scene names in order and the architecture is obvious before you look at a single prompt.
| # | Scene | Narrated | Reference asset |
|---|---|---|---|
| 1 | The Big Bite | yes | The Boy |
| 2 | Teeth Crushing | yes | — |
| 3 | Saliva Mixing | yes | — |
| 4 | Down the Esophagus | yes | — |
| 5 | Into the Stomach | yes | — |
| 6 | Stomach Churning | no | — |
| 7 | Small Intestine | yes | — |
| 8 | Nutrients Absorb | yes | — |
| 9 | Large Intestine | yes | — |
| 10 | The Exit | yes | — |
| 11 | Thumbs-Up Outro | yes | The Boy |
Scene one and scene eleven are the bookends: a live character in a real kitchen, establishing who this is happening to and then closing the loop. Scenes two through ten are the journey, and every one of them is an interior shot with no human in it. That split is the whole design. The character exists to give the viewer someone to care about for four seconds at each end; the middle nine scenes can be pure diagram because the caring has already been established.
The other structural choice worth stealing is scene six. It is the only scene in the recipe with no voiceover at all — the stomach churning, on its own, with nothing said over it. In a 44-second explainer that is a deliberate beat: one shot that is allowed to just be a visual, sitting at the midpoint. Explainers that narrate continuously from first frame to last are exhausting to watch, and the fix is exactly this, one silent scene where the picture carries it.
Where the continuity actually comes from
The recipe has one reference asset: "The Boy," a character with a single reference image, described in enough detail to be reproducible — messy light brown hair, blue eyes, freckles, yellow t-shirt with a red star, blue denim jeans.
Here is the part that surprises people. That asset is referenced by exactly two of the eleven scenes. Nine scenes use no reference image at all.
Which means the visual coherence of this video is not coming from the asset system. It is coming from the prompts. Every single image prompt in the recipe opens with the same style declaration, "3D animated cartoon," and the early scenes extend it to "Pixar-quality rendering." The interior scenes then repeat a consistent vocabulary: cutaway view, soft pink tubes, bright vibrant colors, kid-friendly educational cartoon aesthetic. The look holds across nine unreferenced scenes because nine prompts were written to describe the same look.
This is the right division of labour and it is worth being explicit about it. Use a reference asset when a specific thing must be identical across scenes — a face, a product, a logo. Use repeated prompt language when a style must be consistent but the subjects are all different. Pinning a character reference to a shot of an intestine would achieve nothing, and it would cost you a reference slot in a scene that does not need one. When you do need the asset path, setting up reusable characters and products covers registering one properly.
Narration pacing, scene by scene
Ten scenes carry voiceover, totalling 110 words across 40 narrated seconds. That averages about 2.75 words per second, which is a brisk but comfortable read for a kids' explainer.
Averages hide the interesting part. Per scene, the rate swings hard:
- Scene 9, "Large Intestine": 8 words in 4 seconds. Two words a second, spacious.
- Scene 11, "Thumbs-Up Outro": 7 words in 4 seconds. The slowest line in the video, and correctly so — it is the payoff line and it needs room to land.
- Scene 8, "Nutrients Absorb": 19 words in 4 seconds. Nearly five words a second.
Scene eight is the one to look at if you fork this. It carries the most information in the video (vitamins, energy, blood, powering your body) in the same four seconds every other scene gets, and it is rushed at the exact moment the explainer delivers its actual educational payload. The fix is not a faster read. It is a longer duration for that scene, or a split into two. Uniform durations applied to non-uniform information density is the most common pacing fault in scene-based explainers: every scene here is four seconds, which is tidy, and tidy is not the same as right.
Read your own scene list this way before you render anything. Word count divided by duration, per scene.
Where the credits go
At 396 credits across eleven scenes, this recipe works out to 36 credits a scene at preview resolution and 45 at full. Each of those scenes is two generations under the hood, a still image and then that still animated by the video model, plus the final stitch.
The more interesting number is the gap between the two figures. Most two-figure recipes in the library roughly double from preview to full. Sweet Switch Fruit Heist goes from 441 to 875; Cat-Faced Cabin Heist goes from 496 to 992. This one goes from 396 to 495, a gap of about a quarter rather than double.
That changes the calculus. On a recipe that doubles, the preview pass is an obvious budget lever. Here it saves comparatively little, so the validation you get from it needs to be worth about a fifth of the full render. It usually still is, since continuity and motion are worth checking before you commit, but it is a closer call than the library average.
One thing the 396 does not include: captions. The recipe's tag list mentions a caption look, but the recipe itself does not render captions as part of the run. If your explainer needs burned-in text, adding captions is a separate finishing step on its own pricing, and it belongs on top of your budget rather than inside it.
Forking it for a different subject
The recipe generalises to any staged process: how a payment clears, how a water treatment plant works, how a seed becomes a tree, how a support ticket gets resolved. The steps:
- Write the stages first, as a plain list. The digestive version has nine. Yours will have a different number, and the number should come from the subject rather than from matching this recipe.
- Add the bookends. One opening scene with a person doing the everyday thing that triggers the process, one closing scene with the same person and a one-line payoff. Register that person as a character asset so the first and last scenes match.
- Write one style sentence and put it at the front of every image prompt. The same sentence, verbatim, in every scene. This is what holds the look together across scenes with no reference image.
- Pick your silent scene. One stage in the middle strong enough to run without narration. It is a pacing device, and skipping it makes the whole thing feel relentless.
- Balance words against durations. Count the words per scene, divide by the duration, and fix anything above roughly 2.5 words a second by lengthening the scene or splitting it, before you render.
- Read the whole voiceover aloud as one script. If it does not work as a spoken paragraph, eleven beautiful renders will not save it.
The vertical framing and 44-second runtime here are already sized for where this kind of content gets watched, which is why the shape transfers cleanly to schools and education work. If your subject is a story rather than a process, the AI movie maker surface is the multi-scene path for that instead.
FAQ
How long is the finished video and what format is it in?
Eleven scenes at four seconds each, so 44 seconds of finished video, rendered 9:16 for vertical placements. Every scene is animated with Kling 3 Turbo Image to Video before the clips are stitched together.
Why does the recipe only use its character reference in two scenes?
Because only two scenes contain the character. The other nine are interior cutaway shots with no human in them, and their consistency comes from repeated style language at the front of every image prompt rather than from a reference image. Reference assets pin specific things; prompt preambles pin style.
What is the narration pacing target?
The recipe averages about 2.75 words per second across its narrated scenes, ranging from 1.75 to 4.75. The scene at 4.75 is the one worth fixing in a fork, by giving it more time or splitting it, because that is where the explainer's actual payload lands.
Can I change the number of scenes when I fork it?
Yes, and you usually should, because your subject will not have exactly nine stages. Once you change the count the published total stops applying: work from the per-scene rate of 36 credits at preview resolution instead. Note that the 396 covers generating the scenes and stitching them, not captions or anything else added afterwards, and that swapping the video model changes the per-scene rate for every scene at once.