Sora 2 Pro Storyboard Mode: What Replaced It
Sora 2 is discontinued as of April 2026. How its Storyboard mode chained scenes and held character identity, and what to use instead — VEO 3.1 chaining, Kling 3.0 identity lock — for multi-scene AI video now.
The first time I ran a six-scene short through Sora 2 Pro's Storyboard mode, the lead actress carried her own jawline, hair part, and tired-eye squint across two minutes of cut footage. That sounded like a small thing. It was not. Before Storyboard mode shipped on Sora 2 Pro, doing that meant first-last-frame chaining in VEO 3.1, identity LoRAs in Wan 2.7, or paying a real actor. For a few months, you could write a shot list, hand Sora the reference, and walk away while it cut the picture — until OpenAI discontinued Sora entirely.
This is the practitioner's guide to Storyboard mode as it existed in May 2026 — what it was, how it differed from single-shot Sora 2, the prompt structure each scene wanted, where the duration cliffs were, and when first-last-frame chaining in VEO was still the better answer. Read it now as a record of how the feature worked and a map to its replacements, not a live tutorial.
Updated August 2026: Sora is discontinued. OpenAI pulled the Sora app and web experience — the surface Storyboard mode lived on — on 2026-04-26, and the Sora 2 API shuts down entirely on 2026-09-24. Reported cause: the product was burning roughly $1M/day in compute against about $2.1M in lifetime revenue, and Disney exited the partnership. Sora survives only as a rate-limited generation feature inside ChatGPT Plus/Pro; there's no confirmation Storyboard's multi-scene, multi-reference workflow described below still functions in that reduced form, so treat everything past this point as documentation of a discontinued feature, not a how-to. For multi-scene identity continuity today, the closest replacements are VEO 3.1's first-last-frame chaining (covered in the comparison table below, and it's the one that survives this update intact) and Kling 3.0's start-end frame mode with reference-image identity lock, which held character consistency across scenes better than Sora's token approach even before the shutdown. See the creator survival guide for the full migration picture, and the API sunset migration plan if you have a pipeline still calling Sora directly before September 24.
What Sora 2 Pro Storyboard mode actually was
OpenAI shipped Sora 2 Pro in November 2025. Storyboard mode arrived as a free upgrade in February 2026 and became the default interface on Pro tier — until OpenAI discontinued the whole Sora app and web experience on 2026-04-26. While it worked, it was a multi-scene container that ingested an ordered list of scene prompts plus a shared reference pack (1–6 images of recurring subjects, plus optional style references) and rendered the scenes as one connected piece with a shared identity manifold and color memory.
Three things made it structurally different from running six single-shot Sora 2 generations and stitching:
- Shared identity tokens. A character described once at the top was bound to a token that persisted across every scene. Eye color, scar, freckles, hair part — all sticky.
- Cross-scene color and lighting memory. The grade set in scene one carried into scene six unless overridden. The model treated the storyboard as a sequence shot inside the same DP's coverage.
- Anchored continuity beats. You could mark "Anna picks up the phone" in scene three, and scene four's opening would accept "Anna, still holding the phone, walks to the window" without redescribing the prop.
Single-shot Sora 2 did none of this. Each generation was a fresh world, and getting Anna to look like Anna twice was a coin flip. Today, the closest equivalent to shared-identity multi-scene chaining is VEO 3.1's first-last-frame chaining or Kling 3.0's reference-locked start-end frame mode — see the comparison further down, which is the one section of this guide still describing a live workflow.
How Storyboard mode differed from single-shot Sora 2
Side-by-side, the practical differences while both existed — kept as a reference for anyone evaluating whether a current model's multi-scene mode measures up:
| Capability | Sora 2 single-shot | Sora 2 Pro Storyboard |
|---|---|---|
| Max duration per generation | 20 s | 12 s per scene, up to 12 scenes |
| Total output (one job) | 20 s | ~144 s continuous |
| Character identity across shots | Reference image, fragile after 2 shots | Identity token, stable across 12 scenes |
| Color/grade continuity | None | Inherited unless overridden |
| Audio | Generated per shot, often desyncs across cuts | Co-conditioned across scenes, music bed continuous |
| Reference images | 1 | Up to 6 (subjects + style + props) |
| Cost (Pro tier, retail) | ~$0.12/sec | ~$0.18/sec |
| Failure mode if a scene misses | Regenerate the whole 20 s | Regenerate just that scene, identity preserved |
The last row was the underrated one. Single-shot Sora forced you to throw out 20 seconds of perfectly good footage because second 18 cracked. Storyboard regenerated one 8-second beat and slotted it back in. Kling 3.0's per-scene regeneration on start-end frame mode is the closest thing to this workflow available today.
Scene chaining with reference images: how it worked
Storyboard took a reference pack of up to six images. The pack composition that worked best in practice:
- Subject A — three references: front portrait, three-quarter, full body
- Subject B (if any) — one or two references
- Style/look — one reference still: a frame from a film whose grade and lens feel you wanted to inherit (or a Flux-generated still tuned to match)
You uploaded the pack once at the storyboard level, named each subject (anna, marcus, look_ref), and referenced them by name in scene prompts. The model bound each named reference to an identity token at the start of the job and re-anchored at the start of every scene.
The counterintuitive lesson from running this: don't include too many full-body references. Three was the sweet spot. Five front-on portraits collapsed identity into a single rigid pose and the character couldn't turn their head naturally — a lesson worth carrying into whichever current model's reference-image mode you're using, since the same collapse shows up on Kling 3.0 and VEO 3.1 reference inputs too.
The per-scene prompt structure that wins
Inside a Storyboard job, each scene takes its own prompt. The structure that produces clean output, scene after scene, is the same five-slot template across all twelve scenes — what changes is the content, not the shape.
[Subject reference] + Action + Environment + Camera + Audio
Keep it tight. The model already knows the subject from the reference pack. You do not need to redescribe Anna's hair color in every scene — the token does that. What you must give it every time is the action, the room, the camera move, and the audio bed.
Worked example — six-scene narrative
Reference pack: anna (three stills), look_ref (one moody-blue night still).
Scene 1 (10 s): anna, standing alone on a rain-soaked Brooklyn rooftop at 2am, wide shot from the back at chest height, slow lateral dolly left revealing the skyline, wind audio and distant traffic, no dialogue. Match look_ref grade.
Scene 2 (8 s): anna, turning toward camera as her phone rings, medium close-up, handheld, push-in 0.5 m, phone vibration audio, she answers and says "I'm here." Single line, no other dialogue.
Scene 3 (12 s): anna, descending a fire escape, low angle tracking from below, rain heavier now, metallic clangs of her boots on the grating, no dialogue.
Scene 4 (10 s): anna, stepping into a yellow-lit corner bodega, wide shot from inside the store, fluorescent buzz, distant news radio, she scans the aisle and walks past camera left to right.
Scene 5 (8 s): anna, paying at the counter, over-the-shoulder shot from behind the cashier, warm tungsten light replacing the cool exterior, soft register beep, she says "keep the change."
Scene 6 (12 s): anna, walking out into the rain again, pulling her hood up, wide shot from across the street, the bodega sign reflecting in a puddle, music bed swells gently, fade to black.
Total runtime: 60 seconds. On my second-pass generation this delivered identity intact, grade migrating naturally from blue exterior to tungsten interior and back, and the music bed crossfaded across the scene cuts without a stitch artifact. The only beat I regenerated was scene 4 — first pass had a phantom shopper appear in frame.
Character consistency across scenes — the real numbers
In testing across roughly 30 multi-scene jobs in the last 90 days:
- Scenes 1–6: identity holds nearly perfectly. Subtle differences in micro-expression but the person is unambiguously the same.
- Scenes 7–9: small drift starts. Eye spacing or jaw width can shift 2–4%.
- Scenes 10–12: drift is visible if you A/B scene 12 against scene 1. Not catastrophic — looks like the same actor on a different shoot day — but a careful viewer will notice.
The fix that has worked: every six scenes, refresh the reference pack mid-job by pulling a clean still from your best generated scene and adding it to the pack. The Storyboard interface allows reference editing between scenes. This stabilizes identity past scene 8 substantially.
Duration limits and the cost cliff (as they stood)
Three caps to internalize when reading this as a record of how Storyboard worked:
- 12 seconds maximum per scene. Above that, intra-scene coherence wobbled even on Pro tier.
- 12 scenes maximum per storyboard. Above that, shared identity manifold degraded.
- Total job runtime ceiling: ~144 seconds. Practical ceiling for usable output sat closer to 90 seconds before identity drift became obvious.
The cost cliff lived at the audio layer. Music-bed continuity was co-generated across the whole storyboard, at roughly 1.4x what single-shot generation cost per second.
Storyboard vs VEO 3.1 first-last-frame chaining — and what to use now that Storyboard is gone
Both approaches solved "make a multi-scene piece with the same character." They were not interchangeable, and now only one side of this comparison is still buyable.
| Dimension | Sora 2 Pro Storyboard (discontinued) | VEO 3.1 first-last-frame chaining |
|---|---|---|
| How identity is preserved | Shared identity token, native | Last frame of clip N becomes first frame of clip N+1 |
| Scene transitions | Soft cuts inside the model's color memory | Hard cuts unless you blend in post |
| Dialogue per scene | One short line per scene worked | Strong — VEO is the dialogue king |
| Lip-sync quality | Decent but consonants drifted | Phoneme-accurate in 8 languages |
| Maximum useful runtime | ~90 s clean | ~60 s clean (drift compounds) |
| Camera continuity | Inherited implicitly | You must specify per shot |
| Best for | Narrative, music-driven, multi-location | Dialogue-heavy, single-location, conversation scenes |
Sora 2 Pro Storyboard is no longer available at any price — the app it lived on shut down 2026-04-26 and the API sunsets 2026-09-24. For dialogue-heavy, single- or two-location pieces, VEO 3.1 chaining is now simply the answer, not one side of a trade-off. For music-driven pieces that move through several locations — Storyboard's other lane — the closest current fit is Kling 3.0's start-end frame mode with a locked reference image, which held character identity across scenes at least as well as Storyboard's token approach even before the shutdown, though transitions default to harder cuts than Storyboard's color-memory blending. For a deeper head-to-head on the underlying models, the Sora 2 vs VEO 3.1 deep capability comparison is the reference post.
Exporting as a continuous narrative
Storyboard jobs export three ways:
- Single concatenated MP4 — one continuous file with the model's own scene cuts baked in. Default and the right choice for 80% of work.
- Per-scene MP4s — twelve files, useful if you want to recut in Premiere or DaVinci or layer additional VFX shot-by-shot.
- Edit-friendly bundle — concatenated MP4 plus an XML/EDL with cut points marked, plus per-scene reference frames. This is the one to grab if a real editor is taking the project to finish.
For longer-form narrative beyond Storyboard's 90-second clean ceiling, the typical workflow is: stitch two Storyboard jobs (sharing the same reference pack) and either accept the small mid-piece identity reset as a deliberate "the next morning" cut, or run a final identity-locking pass through the AI movie maker which handles cross-job identity bridging.
Common failure modes and the fixes
- Scene 7 onward identity drift. Refresh reference pack at scene 6 with a clean still pulled from scene 4 or 5.
- Music bed clips at scene cut. Add "music bed continues across cut" in the next scene's prompt, or specify the same instrumentation across both scenes.
- Wardrobe wandering. Sora's identity token does not bind wardrobe as tightly as face. Restate the wardrobe in scenes that introduce a new location.
- Phantom characters appearing in wide shots. Add "no other people visible" inline. Storyboard respects negatives the same way single-shot does.
- Soft cuts feel too soft. If you want a hard cut between scenes, write "hard cut to:" at the start of the next scene's prompt. The model honors it.
- Dialogue paraphrased. Same fix as single-shot — quote the line, keep under 12 words, and for longer dialogue regenerate silent and run a lip-sync pass on top.
Where Storyboard used to fit in a real production
A typical 90-second narrative piece ran through this process while Storyboard was live:
- Write the shot list — 6 to 10 scenes, with the same five-slot prompt structure for each.
- Build the reference pack — three stills of each recurring character (generated in Flux or Nano Banana if you don't have a real shoot).
- Run Storyboard end-to-end. Total wall time at Pro priority ran 12–20 minutes.
- Identify the one or two scenes that missed. Regenerate them at the scene level with tightened prompts.
- Drop into the edit. If dialogue clarity mattered, replace VO with voice cloning and run lip-sync.
- Color match any Storyboard-internal grade jumps in DaVinci or directly in AI movie maker.
That process no longer has a Storyboard step to run. If you're working from a written narrative, story to video still handles the shot breakdown and routes each beat to whichever current model fits it — VEO 3.1 for dialogue close-ups and first-last-frame chaining, Kling 3.0 for the music-driven or multi-location connective spine, Lyria or Suno for the score. The steps above map cleanly onto that roster; only the model in step 3 changed.
For broader context on which video model wins which shot type, the best AI video generation models 2026 breakdown remains the right reference, and the creator survival guide covers the full migration picture.
FAQ
How was Sora 2 Pro Storyboard different from regular Sora 2?
Storyboard took a multi-scene shot list plus a shared reference pack and rendered all scenes as one connected piece with shared identity tokens and color memory. Regular Sora 2 was single-shot — each generation was a fresh world. Storyboard was what you reached for when the same character had to appear in scenes 1, 4, and 9. Neither is purchasable anymore — OpenAI discontinued the Sora app on 2026-04-26 and the API follows on 2026-09-24.
What was the maximum length of a Sora 2 Pro Storyboard job?
12 scenes, 12 seconds each, ~144 seconds total. Practical clean-output ceiling sat around 90 seconds before identity drift became visible to a careful viewer. For current multi-scene work, VEO 3.1's chaining runs about 60 seconds clean; Kling 3.0's extended mode reaches roughly 30 seconds per clip.
Did Sora 2 Pro Storyboard handle dialogue and lip-sync?
It generated dialogue audio across scenes with the same voice timbre, but lip-sync was acceptable, not phoneme-accurate. VEO 3.1 still wins on dialogue today, phoneme-accurate lip-sync included.
Can I use my own actor's face as a reference on a current model?
On VEO 3.1's reference-to-video and Kling 3.0's reference-image modes, yes — upload stills of the actor as the subject reference. License-wise, you must have rights to use that likeness, and both providers' content policies reject celebrity faces and public figures the same way Sora's did.
Storyboard or VEO first-last-frame for narrative pieces — which do I use now?
Storyboard isn't an option anymore. VEO 3.1 chaining is the default answer for dialogue-heavy, one- or two-location pieces. For pieces that move through several locations with music carrying the cuts — Storyboard's other lane — Kling 3.0's start-end frame mode with a locked reference image is the closest current fit.
Is Sora 2 Pro Storyboard still available on Versely?
No. Sora 2 Pro Storyboard was tied to the Sora app and API, and OpenAI discontinued the app on 2026-04-26 with the API following 2026-09-24. Versely's AI video generator routes multi-scene narrative work to VEO 3.1 (dialogue, first-last-frame chaining) and Kling 3.0 (music-driven, multi-location, reference-locked identity) instead, assembled in the AI movie maker and story to video tools.
Bottom line
Sora 2 Pro Storyboard was one of the first AI video tools to treat narrative as a first-class object — continuity in the generation instead of faked in the edit — but OpenAI discontinued it along with the rest of Sora: the app on 2026-04-26, the API on 2026-09-24. The lesson it leaves behind still holds: route dialogue beats to VEO 3.1, the connective multi-location spine to Kling 3.0's reference-locked chaining, and the score to Lyria or Suno, assembled through Versely's router — a one-day creative director can still ship a 90-second narrative piece that holds up next to footage that took a week, just without Sora in the stack.