Wan 3.0 Just Topped the Video Leaderboard. The Open-Weights Story Is Now a Video Story.

For most of the generative AI boom, the open-weights conversation belonged to images. Stable Diffusion proved you could run a real model on your own GPU, and a whole ecosystem of LoRAs and ControlNets grew around it. Video was different. The best video models stayed locked behind APIs, and "open source video" meant small models that lagged the closed leaders.
That changed this month. Wan 3.0, Alibaba's video model, took the top spot on Artificial Analysis's next-generation AA-Video-T2V v2.0 benchmark, built from more than 68,000 human preference votes across roughly a thousand prompts. Wan 3.0 posted an Elo of 1157 and won first place in 10 of 20 category rankings. The same leaderboard put Wan 3.0 first in video editing too, ahead of ByteDance's Seedance and MiniMax's H3.
The significance is less about one number and more about what it represents. An open-weights Chinese model is now competitive with, and in several rankings ahead of, Google's Veo 3.1 and the rest of the closed field. The pricing tells the same story. Wan 3.0 runs $0.05 to $0.20 per second depending on resolution, against Veo 3.1's $0.40 at standard tier. Alibaba raised $10 billion in a Hong Kong share placement to fund exactly this push, and the strategy is legible: sell cheap, gain share, make the enterprise customer the growth engine.
Technically, Wan 3.0 is built for consistency and structured input. It generates clips up to 30 seconds in a single pass, doubling its predecessor's 15-second limit, at up to 1080p with synchronized micro-expressions and multilingual voice. More interesting, it accepts documents, PDFs, PowerPoints, and spreadsheets as input, turning a slide deck or a report into video. That document-to-video angle is aimed squarely at enterprise workflows, not just individual creators.
The document-to-video feature deserves more attention than it gets. Most video models want a text prompt or an image. Wan 3.0 wants your existing assets, the deck you already made, the report you already wrote, and it treats those as the source of truth. For a business that lives in documents, that lowers the friction of going from "we have a report" to "we have a video" to almost nothing. It is the first model to make the enterprise document pipeline the on-ramp, and it is a quietly smart move.
For a team choosing a video stack, this reshuffles the decision. The old assumption, that open models are what you reach for when you cannot afford or cannot use the closed ones, is weakening. You can now self-host a model that tops the leaderboard, which matters for the cases that have nothing to do with quality and everything to do with control. A client whose data cannot leave their infrastructure has no compliant option other than a self-hosted model, and for the first time that option is not a compromise on output quality.
There is a caveat that gets skipped in the enthusiasm. Self-hosting is not free just because the weights are. You pay in GPUs, in integration work, in the ongoing job of keeping a pipeline healthy at whatever volume your work actually reaches. The model is free; the engineering is not. For most small teams, a hosted endpoint still wins on total cost until volume justifies the hardware.
The benchmark context matters for reading the win correctly. The AA-Video-T2V v2.0 benchmark is the new generation of human-preference evaluation, and it judges every model at 1080p, which removes the old trick of winning on low resolution. Wan 3.0 took first in ten of twenty category rankings and posted the top overall Elo at 1157. That is not a niche win or a favorable category; it is a broad, competitive result against the entire field, including the models that cost eight times as much per second. When an open model that costs $0.05 a second beats a closed model that costs $0.40, the value proposition is hard to argue with.
There is also a political dimension that is hard to ignore. The AI video race has become a proxy for a larger competition between Chinese and American labs, and Wan 3.0's benchmark crown is the kind of result that gets cited in both countries for different reasons. In China it is proof the domestic stack can compete on quality, not just price. In the US it is evidence that the cost of compute is not the only thing shifting. Neither framing fully captures what is happening, which is a genuine competitive convergence that benefits users in the form of lower prices and more capable open models.
The open-weights crossover into video also has implications for the broader geopolitics of AI. Chinese labs have been winning the price war for a while, and now they are winning on benchmarks too, which changes the narrative from "cheap imitation" to "genuinely competitive." The same dynamic already played out in image generation with Qwen-Image and Z-Image, and in language models with DeepSeek and Qwen. Video was the last major holdout for the closed Western labs, and Wan 3.0 just showed the door is open.
The benchmark crown is the headline, but the real story is that video generation now has a credible open road. The same dynamics that followed Stable Diffusion are likely to follow here: fine-tunes, community control nodes, regional deployments, and a long tail of use cases the API vendors never bothered to serve. Fine-tuned open models are already the standard way to hold a specific character or style in image work, and video creators will want the same ability to train on their own brand or product.
There is already early evidence of that ecosystem forming. The MiniMax H3 ecosystem has produced an open-source prompt builder and a wave of community fine-tunes aimed at anime and character work, mirroring the Stable Diffusion playbook. When the open weights are good enough to top a benchmark, the community that forms around them is no longer a niche; it is the mainstream, and the innovations that come out of it, control nodes, regional hosting, privacy-preserving local inference, tend to arrive faster than any single vendor can ship.
The practical implication for a team is concrete and worth restating. If you are choosing a video stack in late 2026, the question is no longer "open or closed" as a matter of quality. It is a matter of control, cost, and exit risk. Open weights give you the first and reduce the third, at the cost of the second. Closed APIs give you convenience at the cost of all three. The teams that will navigate the next model deprecation without pain are the ones who already decided where they stand on that tradeoff, instead of finding out the hard way when their vendor announces a shutdown.
For anyone building on generative video, the practical lesson is to stop treating open weights as a fallback. Treat them as a strategic option. The next time a closed vendor deprecates a model the way OpenAI just did with Sora, the teams that are not scrambling will be the ones who kept at least one foot on the open road. Open weights used to be the consolation prize. In video, they just became the trophy case.
Related articles
A Wheeled Semi-Humanoid Finished an Hour of Laundry Without Help
Individual tasks can succeed while a workflow still fails. Dyna changed the metric.
LTX 2.5 Wants to Render Your Blocky Blender Draft Into a Finished Shot
You do not control what happens in text-to-video. This tries to fix that.
ServiceNow Turns Agent Failures Into Training Data
Generation without verification is noise. The gates are the product.
One Framework for Language and Vision: Horizon's 1.6B Open Model
A bet that the bridges between language and vision were never needed.