← Back to glossary
+Suggest a term
Tool·AI Models & Capabilities·Added 16 days ago

Wan3.0

Also known as: Wan 3.0, wan3.0-video, Alibaba Wan

Alibaba's video generation model, launched in public beta August 2026. It generates up to 30 seconds of video in a single pass from text, images, audio, video, and for the first time, documents like PDFs and slide decks.

Wan3.0 is Alibaba Cloud's third-generation video generation model, entering public beta on August 6, 2026 via Alibaba Cloud Model Studio and Qwen Cloud. The headline capability is native 30-second single-pass video — double the 15-second ceiling of its predecessor — which means continuous camera movement and unbroken shots without manually stitching clips together.

The model's most distinctive addition is document input. Alongside text, images, audio, and video, Wan3.0 accepts PDFs, spreadsheets, slide decks, and web pages as creative references, effectively turning a PowerPoint into a video brief. Alibaba calls this Omni-Reference. API pricing starts at $0.05 per second for 480p output, scaling to $0.20 per second for 1080p, so a full 30-second 1080p clip costs roughly $6.

As of launch, Wan3.0 has no published weights, no model card, and no independently verified benchmarks. Full API access was described as coming 'soon.' For builders evaluating it: treat the pricing tiers as the reliable signal, ignore unverified claims about 4K support or parameter counts circulating on social media.

This definition is AI-generated and refreshed weekly. It may contain inaccuracies. Use your own judgment, especially for production decisions.
Related terms
Text-to-videoDocument-to-videoMultimodal modelSeedanceKling