AI News

    Wan 3.0 Beta: 30-Second Takes and Document-to-Video

    Alibaba's Wan 3.0 beta claims 30-second single takes and video generated from PDFs, decks, and spreadsheets. A skeptical look, and what Wan 2.7 does today.

    Versely Team7 min read

    Alibaba opened a public beta for Wan 3.0 in early August, and the two headline claims are the kind that normally take a full model generation to arrive separately: 30-second single-take clips, and video generated directly from documents — PDFs, slide decks, spreadsheets, even a plain URL. Coverage of the beta has been consistent enough across independent outlets that the core claims look real. Whether they hold up outside curated beta demos is a different question, and "public beta" is doing a lot of work in that sentence.

    Open laptop displaying a slide deck presentation next to a notebook

    What Alibaba actually announced

    According to TNGlobal's report on the public beta launch, Wan 3.0 supports clips of up to 30 seconds in a single generation pass — double the 15-second ceiling of Wan 2.7 — and accepts multimodal inputs including text, image, video, audio, web pages, PDFs, and PowerPoint presentations. A second independent write-up, covering the same release, confirms the document list in more detail: pdf, doc, xls, ppt, and md files, plus web pages by URL, feeding directly into video generation.

    Put together, the pitch is that a slide deck, a spreadsheet, or even a webpage can become a starting point for a video the same way a text prompt or image currently does — no separate scripting or storyboard step required before generation begins.

    The skeptical beat

    A few things are worth holding onto before treating this as settled fact. First, this is a public beta, and beta access has historically meant gated, rate-limited, or region-restricted availability rather than a full production rollout — the coverage confirms the beta is live and testable through Alibaba's own platforms, not that it's generally available at production quality and volume. Second, the document-to-video claim is a genuinely hard problem: turning a dense slide deck into coherent, well-paced video requires the model to make a lot of editorial decisions — what to show, for how long, in what order — that a text-to-video prompt never has to make. Early beta output on that specific capability is the part most likely to need another few iterations before it's reliable for real production work.

    None of that means the underlying claims are false. Two separately reported sources landing on the same 30-second figure and largely overlapping document-type lists is a reasonable bar for "this is real," even without hands-on testing. It just means "beta" should be read literally.

    What document-to-video could mean for explainer creators

    If the capability matures the way it's described, the most obvious win is for exactly the kind of content that currently requires the most manual scripting: product explainers, onboarding walkthroughs, and quarterly-update-style videos built from material that already exists as a document. A ten-slide deck that already states the problem, the solution, and three feature callouts is most of a video script already written — the labor today is translating that into shot-by-shot prompts. A model that can ingest the deck directly and propose a first-pass video collapses that translation step, even if a human still needs to review pacing and trim the result afterward.

    That's the theory. It's not something Versely runs today, because Wan 3.0 isn't stable or broadly available yet — but it's worth tracking if your content pipeline leans on decks, one-pagers, or written briefs as source material, which is exactly the audience for tools like Versely's AI explainer video generator and story-to-video.

    What Wan 2.7 already does on Versely, right now

    While Wan 3.0 stabilizes, the current generation — Wan 2.7 — is a genuinely capable, production-ready model that's already running on Versely across three endpoints. Wan 2.7 was itself a notable step up from earlier Wan releases: it introduced first/last-frame control, where you set the opening and closing frame of a shot and the model fills in the motion between them, alongside a 15-second clip ceiling that Wan 3.0's 30-second claim is now trying to double. That first/last-frame feature is worth knowing about on its own — it's a different kind of control than reference-locking, closer to directing the exact start and end pose of a shot than to holding a character's identity steady.

    Wan 2.7 Text to Video generates from a text prompt alone, with clips up to 1080p and duration from 2 to 15 seconds. It's billed 10 credits per second at SD/HD and 15 credits per second at 4K on Versely.

    Wan 2.7 Reference to Video takes up to five reference images or videos plus an optional first frame and voice-timbre control, generating 2-to-10-second clips at 12 credits per second SD/HD or 18 per second 4K. This is the closest thing Wan 2.7 has to Grok Imagine's multi-reference approach — useful for holding a product or character steady across a short sequence today, without waiting on 3.0.

    Wan 2.7 Video Edit modifies an existing source video using a text instruction and an optional reference image for character, clothing, or style guidance, at the same 12/18 credits-per-second tiers, for 2-to-10-second edits.

    A working walkthrough with what's live today

    You don't need to wait for document-to-video to compress a deck into a first video pass. A practical version of the same idea, buildable on Versely right now:

    1. Pull the three or four key claims from your source document manually — the headline stat, the problem, the solution, the CTA.
    2. Generate a reference image for your product or subject, then run it through Wan 2.7 Reference to Video with a short scene prompt per claim, keeping the reference locked across all of them so the product stays visually consistent scene to scene.
    3. Use Wan 2.7 Video Edit on any individual scene that needs a targeted fix — a color shift, a different camera angle — instead of regenerating the whole clip from scratch.

    It's a manual version of what document-to-video promises to automate. The gap Wan 3.0 is trying to close is the manual extraction step in point one — everything after that already works.

    FAQ

    Has Wan 3.0 fully launched? It's in public beta as of early August 2026, according to independent coverage of Alibaba's release. Public beta typically means gated or limited access rather than a full general-availability rollout.

    What file types does Wan 3.0 claim to accept for document-to-video? Reported coverage lists PDF, DOC, XLS, PPT, and MD files, plus web pages supplied by URL.

    How long can a single Wan 3.0 clip run? Up to 30 seconds in one generation pass, per beta coverage — double Wan 2.7's 15-second ceiling.

    Does Versely support Wan 3.0 yet? Not yet — Versely currently runs Wan 2.7 across Text to Video, Reference to Video, and Video Edit. Wan 3.0 will be evaluated as the beta stabilizes.

    Is Wan 2.7 worth using now, or should I wait for 3.0? Wan 2.7 is stable, production-ready, and already covers text-to-video, reference-to-video, and video editing at up to 1080p. Waiting on a beta for a feature you can approximate manually today is rarely the better trade for anything you need to ship this month.

    Closing takeaway

    Wan 3.0's beta claims are ambitious and, on the current evidence, real — but "real" and "reliable at scale" are different bars, and public beta software rarely clears the second one on day one. Document-to-video is the more interesting claim of the two, because it targets the actual bottleneck in explainer and update-video production: turning something you already wrote into something you can watch. Until that capability is proven out beyond beta demos, Wan 2.7's reference-to-video and video-edit endpoints on Versely cover most of the same practical ground today, one deliberate step at a time instead of one automated pass.