Wan 3.0 entered public beta on August 6, 2026 through Alibaba's Qwen Cloud and DashScope API, priced per second at $0.05 (480p), $0.10 (720p), and $0.20 (1080p). It suits teams needing 30-second single-shot video or document-to-video ad creation, not teams needing open weights to self-host.
Wan 3.0 is Alibaba Tongyi Lab's video model, in public beta since August 6, 2026, generating native single-shot clips up to 30 seconds at up to 1080p. It uniquely accepts PDFs, slide decks, and spreadsheets as creative references alongside text, image, audio, and video.
Provider: Alibaba Cloud · Family: Wan
More about Alibaba Cloud on HokAI
Input modalities: text, image, audio, video, pdf · Output: video
About Wan 3.0
Wan 3.0 is a video generation and editing model built by Alibaba's Tongyi Lab, the team behind the earlier Wan 2.1 through 2.7 releases. Alibaba Cloud opened public beta access on August 6, 2026, listing the model on its Qwen Cloud and Model Studio platforms as wan3.0-video, reachable through the DashScope international API endpoint. It sits at the top of the Wan lineage as the first release since Wan 2.2 to drop the numbered-point branding pattern, following a stretch of adjacent Tongyi releases (Wan-Dancer-14B, WanSong, Wan-Streamer) in July 2026 that expanded the family into music, dance, and real-time streaming before the flagship video model itself shipped. The defining capability is native single-shot generation up to 30 seconds long, roughly double the 15-second ceiling of Wan 2.7 and far past the 5-8 second clips typical of most 2026 video models. Alibaba markets this as removing the need to stitch clips for continuous camera movement or one-take shot language. Output tops out at 1080p; there is no confirmed 4K tier despite that claim circulating on unofficial fan sites. Independent benchmark scores (VBench or similar third-party leaderboards) have not been published for Wan 3.0 as of this writing, so no comparative score can be verified against Veo, Kling, or Seedance. Wan 3.0's standout feature is Omni-Reference input: alongside the usual text, image, audio, and video prompts, it accepts documents, spreadsheets, slide decks, PDFs, and live web pages as creative references, turning a slide deck or product spec into a generated video. Alibaba folds reference generation, instruction-based editing, subject replication, and motion-driving into one unified model rather than requiring separate specialized checkpoints, continuing the direction Wan 2.7 started with first/last-frame control and multi-reference input. Pricing on Qwen Cloud runs per second of output at three resolution tiers: $0.05/sec at 480p, $0.10/sec at 720p, and $0.20/sec at 1080p, meaning a full 30-second 1080p clip costs about $6.00. Third-party aggregator AIHubMix lists a lower resale rate for the same three tiers. There is no published flat-rate or subscription plan; access is metered API usage only. Despite widespread rumors that Wan 3.0 shipped as an Apache 2.0 open-weight release in 1.3B and 14B sizes, no weights or inference code for Wan 3.0 itself have been published on the Wan-AI Hugging Face organization, the Wan-Video GitHub organization, or ModelScope as of this writing. That open-weight claim appears to trace to a single unverified LinkedIn post rather than any first-party Alibaba channel, and is likely confusion with the genuinely open Wan 2.1/2.2 checkpoints or the adjacent Wan-Dancer/WanSong releases, which did ship under Apache 2.0. Treat Wan 3.0 itself as closed, API-only, and proprietary until Alibaba states otherwise. Full third-party API access was described by Alibaba as opening "soon" without a committed date at beta launch; the model is currently reachable through Alibaba's own surfaces (Model Studio, the Wan website, the Qwen PC creation tool) and through early aggregator listings like AIHubMix. Regions, SDKs, and enterprise deployment options beyond the DashScope international endpoint have not been detailed publicly. Alibaba has not published a system card, safety benchmark, or training-data disclosure specific to Wan 3.0 at time of writing. As with prior Wan releases, expect a visible or embedded content provenance marker on generated output, though this has not been independently confirmed for the 3.0 beta. Wan 3.0 fits teams that need a single continuous shot longer than a few seconds, document-to-video workflows for ads or product explainers, or unified reference/edit/motion-drive in one call. It is a weaker fit for anyone who needs open weights to self-host or fine-tune (use Wan 2.2 or Wan 2.7 instead), needs independently verified quality benchmarks before committing budget, or needs guaranteed general API availability today rather than a beta waitlist.
Pricing
Qwen Cloud bills per second of generated video: $0.05/sec low tier, $0.10/sec mid tier, $0.20/sec top tier, so a full-length top-tier clip runs about $6.00. Third-party reseller AIHubMix lists a discounted rate near $0.042/$0.085/$0.169 per second across the same three tiers. No flat-rate or subscription plan is published.
Key Features
- Long Single-Shot Generation: Native continuous generation in a single pass, well beyond the short-clip ceiling typical of most 2026 video models, avoiding clip-stitching for continuous camera movement.
- Document-to-Video (Omni-Reference): Accepts PDFs, Word, Excel, PowerPoint files, markdown, and live web pages as creative references alongside text, image, audio, and video, turning a slide deck or spec sheet into a video.
- Unified Reference, Edit, Replicate, and Motion-Drive: One model handles reference generation, instruction-based editing, subject replication, and motion driving instead of requiring separate specialized checkpoints.
- Metered High-Definition Output: Generates across three metered resolution tiers, topping out at full HD; no confirmed ultra-HD tier despite unofficial claims.
Pros
- Native single-shot generation well past the short-clip norm most 2026 video generators impose, cutting out manual stitching between takes.
- Omni-Reference document input (PDF, slides, spreadsheets, web pages) is a differentiator no major competitor publicly matches yet.
- Consolidates reference, edit, replicate, and motion-drive into a single model call instead of separate tools.
Cons
- No independently verified benchmark scores exist yet, so quality claims versus Veo, Kling, or Seedance cannot be checked.
- Closed weights and API-only despite persistent rumors of an Apache 2.0 release; cannot self-host or fine-tune.
- Full third-party API access remains gated at beta launch with no committed general-availability date.
Frequently Asked Questions
How much does Wan 3.0 cost per video?
Wan 3.0 bills per second of generated video on Qwen Cloud: $0.05 at 480p, $0.10 at 720p, and $0.20 at 1080p, so a full 30-second 1080p clip costs about $6.00. Third-party reseller AIHubMix lists a discounted rate near $0.042 to $0.169 per second across the same tiers. There is no published flat-rate or subscription plan.
How does Wan 3.0 compare on benchmarks vs Kling or Veo?
No independently published benchmark scores (VBench or equivalent) exist for Wan 3.0 as of August 2026, so a verified head-to-head against Kling or Google Veo cannot be made yet. Its distinguishing edge is a 30-second native single-shot ceiling, roughly double Wan 2.7's 15 seconds, plus document-to-video input that neither competitor publicly offers.
Is Wan 3.0 open source or proprietary?
Wan 3.0 is proprietary and API-only. Despite widespread claims that it shipped Apache 2.0 weights in 1.3B and 14B sizes around April 2026, no weights or inference code for Wan 3.0 have appeared on the official Wan-AI Hugging Face org, the Wan-Video GitHub org, or ModelScope. That claim traces to a single unverified post, not an Alibaba channel; the genuinely open releases are Wan 2.1 and Wan 2.2.
Does Wan 3.0 train on user data?
Alibaba has not published a data retention or training-on-inputs policy specific to Wan 3.0 as of this writing. No system card or safety disclosure has been released for the model, so this should be confirmed directly with Alibaba Cloud before sending sensitive content.
Who is Wan 3.0 best for and who should avoid it?
Wan 3.0 fits ad and e-commerce teams generating video directly from documents or decks, and editors needing a single continuous shot longer than a few seconds. Teams that need self-hostable or fine-tunable weights should use the genuinely open Wan 2.2 or Wan 2.7 instead, and buyers who need benchmark-verified quality before committing budget should wait for independent scores.