WAN 3.0 — what Alibaba's new video AI model can do

WAN 3.0 — what Alibaba's new video AI model can do

On August 6 Alibaba opened the public beta of Wan 3.0 — the next generation of its Tongyi Wanxiang video model. Two headline claims: clips of up to 30 seconds in a single take, and — an industry first — video generated from documents: the model accepts pdf, doc, xls, ppt files and even links to web pages. Here is what the model can do, with the promo and official demos — and with the facts separated from the rumors already swirling around the release. GPTunneL currently offers WAN 2.7, the line's current API flagship; we'll add 3.0 as soon as Alibaba opens a public API.

Here is the Wan 3.0 promo — 30 seconds of continuous action shot "in one take":

The headline: 30 seconds in one take

Previous versions of the line topped out at 2–15 seconds per generation — Wan 3.0 stretches a continuous clip to 30 seconds in a single pass, with no stitching of short fragments. Alibaba frames the shift as going "from generating a frame to telling a whole story": long camera moves with characters and scenes staying consistent the entire way. Audio is generated in the same pass as the picture, and "smart duration" recommends a clip length to match the prompt — a short scene won't get padded with empty seconds.

Physics and light handling show well in a simple scene — a soap bubble on the beach:

Video from documents — an industry first

The release's signature feature, one no competitor has: besides text, images, audio and video, Wan 3.0 accepts doc, xls, ppt, pdf and md files (and, per Chinese press, also txt, key, pages and numbers) up to 100 MB or 50 pages, plus web pages by URL. From these the model builds a video sequence of up to 30 seconds — the official announcement describes it as transforming "static, text-heavy data into reality-grade video content".

The practical upside is obvious: a training video from a manual, a product demo from a slide deck, a video version of a quarterly report from a spreadsheet — no scriptwriter or editor involved. That's exactly the business angle Alibaba pushes in the announcement.

One model instead of a toolbox

In the 2.x line, references, editing and voice lived in separate models — Wan 3.0 merges everything into one. Alibaba calls it Omni-Reference: the model keeps characters, props, voices and style consistent across shots and tasks.

Editing finished clips is especially interesting: you can select a time interval and regenerate only that part while leaving the rest untouched, and adjust the visuals, the plot and even the characters' lines. Chinese press describes it as moving "from gacha to production" — instead of endlessly rerolling the dice, you get a controllable process.

How confidently the model holds a complex frame shows in a commercial-style clip with on-screen graphics — interfaces and text in frame have always been a weak spot for video models:

Quality: 1080p, micro-expressions — and honest limitations

Resolutions are 480p, 720p and 1080p. The "Wan 3.0 does 4K" claims circulating online are misinformation: Alibaba has no 4K tier and no mention of 4K anywhere in the announcement. The release has spawned a whole wave of fakes — affiliate sites credit the model with open weights, "Apache 2.0" and free access; none of it checks out.

What does check out is the focus on people in frame: detailed faces and skin, restrained natural emotion, micro-expressions tied to gestures. Chinese reviews are honest about the rough edges too: audio quality and text accuracy in frame still have room to grow.

Alibaba shows the cinematic range with genre demos — sci-fi:

And stylized 3D animation on the level of a feature film:

Wan 3.0 against the competition

There are no independent benchmarks for 3.0 yet — the model hasn't appeared in Video Arena. The line's track record is the best guide: WAN 2.7 holds 4th place in the ranking (Elo 1161), behind only Gemini Omni Flash, MiniMax H3 and Seedance 2.0. The direct rival on duration is Seedance 2.5, which got to 30-second one-takes a week earlier; ByteDance is stronger on references (up to 50 inputs), while Alibaba has documents and web pages, which nobody else offers.

One important caveat: Wan 3.0 is a closed model. No weights, no ComfyUI nodes, API only. The line's last open flagship is Wan 2.2 (Apache 2.0, summer 2025), and the community still reminds Alibaba of the broken promise to open-source 2.5.

Where to try Wan 3.0

  • Alibaba's own services — Model Studio (Bailian), the Qianwen app, creative platforms like Wanjing Yike. All built for the Chinese market: Chinese interface, Chinese phone number required to sign up.
  • QwenCloud API — international access, but the rollout is gradual.
  • You can't download Wan 3.0 — the weights are closed; any "download Wan 3.0 for free" site in search results is a fake.
  • In GPTunneL — the Wan line is already in the lab: WAN 2.7, plus the cheaper WAN 2.6 and WAN 2.5. No VPN needed, local payment methods supported. We'll add Wan 3.0 the day the public API opens — just as we did with every model in the line.

How much Wan 3.0 costs

There is no free access: the beta in Alibaba's services is metered, and the international API costs $0.05, $0.10 and $0.20 per second in 480p, 720p and 1080p respectively — a 30-second 1080p clip runs about $6. That's a third more than Wan 2.7, with no independent quality benchmarks in sight yet.

In GPTunneL no subscription is needed — you pay per second of finished video: WAN 2.7 costs from $0.20 per second in 720p and $0.30 in 1080p. Draft on the cheaper WAN 2.5/2.6, render the final cut with the top model; the calculator on the pricing page gives an exact quote for your clip length.

Wan 3.0 FAQ

Is Wan 3.0 free? No. The public beta in Alibaba's services is metered and the API is paid. In GPTunneL the Wan line charges per second of finished video, with no mandatory subscription.

Can I download Wan 3.0? No, the weights are closed — the model is only available through cloud APIs. "Free download" sites in search results are not Wan 3.0.

How is Wan 3.0 different from WAN 2.7? Duration grew from 15 to 30 seconds in one take, documents and web pages were added as inputs, references and editing merged into a single model, and you can now regenerate a selected interval and edit characters' lines.

What languages can Wan 3.0 speak? Alibaba hasn't published an official voice language list for 3.0. The line has a strong multilingual track record though: WAN 2.7 in GPTunneL generates speech in many languages right now.

When will Wan 3.0 arrive in GPTunneL? As soon as Alibaba opens the public API — new models land here the day access opens.

Try it yourself

While Wan 3.0 is on its way to a public API, the rest of the line is ready to go: open WAN 2.7 in the GPTunneL lab — 1080p, multilingual speech, no subscriptions, pay per use. And to catch the 3.0 release the moment it happens, follow the news in our blog and Telegram channel.