Seven Universities Jointly Open-Source OpenWAM: Robots Need More Than "Seeing"—They Need to "Predict the Next Second"

In early October, a team of researchers from seven universities, including the National University of Singapore, Tsinghua University, and Peking University, released OpenWAM on GitHub and Hugging Face. The project's positioning is clear: an open-source research full stack for World Action Models (WAMs), plus a base model. The paper is already on arXiv, numbered 2609.07398, and the repository had about 940 stars by October 1.
The news didn't take long to spread in China, but what truly deserves a closer look is the old problem it aims to solve: it's not enough for a robot to merely recognize what is in front of it.
Traditional simulators compute physics and vision separately
In the past, there were two common paths for robot policies. One is to directly learn visual patterns and control signals simultaneously from robot demonstrations; the other is to first undergo image-text pretraining, bringing in semantic and language knowledge—this is the VLA route.
World Action Models take a third path: first inherit the prior of "how the world will change" from video generation pretraining, then learn "what to do in this world" through embodied experience. It may sound like a detour, but it avoids the old rift between simulation and real robots. Traditional simulators often compute physics and vision separately, so the effect of actions and the visual appearance don't match up; OpenWAM's approach lets the model directly learn "how actions change the world" from data, and then use that to predict object positions, scene semantics, and task states.
Split into three layers to make experiments reproducible
The team breaks the problem into three layers, each corresponding to a component.
OpenWAM-Infra is modular infrastructure that decouples four layers: composable models, training runtime, deployment runtime, and evaluation protocols. The role of this layer is to turn World Action Models from a tightly coupled single implementation into a testbed with swappable parts.
OpenWAM-Study is responsible for distilling design principles. It doesn't build models; it conducts controlled experiments, using sets of comparisons to turn "what kind of design works" into reproducible conclusions.
OpenWAM-α is the base model itself, validating the entire approach on large-scale human first-person videos and robot data.
Three design principles, each backed by comparison numbers
Papers like this fear being all slogans. All three conclusions OpenWAM provides come with controlled experiments.
First, a sufficiently strong generative world prior is needed. A 14B video backbone achieved 93.79%, while 1.3B got only 90.14%; the default configuration chose 5B, just 1.4 percentage points below 14B. Combined with the compact visual latent space compressed by DINOv3 and S-VAE, the performance approaches Wan2.2's built-in VAE.
Second, a clear information flow from world to action must be established during training. Letting the "action stream see the world stream" achieved 92.39% accuracy; switching to isolated masking gave only 87.41%, a full 5 percentage points lower. Synchronous joint denoising at inference works best.
Third, data should be used by source. Robot data is responsible for preserving action grounding, human first-person videos are responsible for expanding world coverage, and single-stage co-training integrates the two. Interestingly, in-domain pretraining improvement is limited, but out-of-distribution (OOD) success rate rose from 14.50% to 26.62%.
Actual specifications of the base model
OpenWAM-α is composed of two parts: Wan2.2-TI2V-5B as the video stream, plus a 1B action DiT attached. The action space is unified into 80 dimensions, compatible with 17 physical platforms. Pretraining data totals 14,300 hours, of which human first-person perspective accounts for 30%, synthetic data 30%, and real data 40%.

In terms of inference speed, a single action chunk takes 170 milliseconds on an RTX 5090. Evaluation covers 8 simulation benchmarks including LIBERO, RoboTwin2.0, and RoboDojo, with embodiments ranging from single-arm and dual-arm to mobile manipulation and dexterous hands. On EBench's mobile dual-arm task, it is currently the best, leading the second place by about 4 percentage points. For real-robot validation, it ran six tasks with a Franka single arm, achieving an average success rate of 82.5%, higher than LingBot-VA's 77.5% and π0.5's 55.0%, with two tasks reaching 100%. Switched to a dexterous hand unseen during training, after fine-tuning it also comprehensively surpasses π0.5.
It does not declare VLA out of the game
One conclusion in the paper is worth pulling out separately: WAM uses future video latent variables for supervision and fits better in-distribution; VLA is more robust under out-of-distribution perturbations. The contest between the two is not settled.
This statement is far more honest than "a certain paradigm is dead." Readers of papers easily remember only the scores, but what they should remember more is this: the two routes now each have their strengths, and it is not yet time to make a final call.
The barrier has been lowered a notch
Viewed against the larger backdrop, on the world model line, Google's Gemini Robotics 2 proclaimed "one brain, applicable to any robot," while Nvidia's Jim Fan threw out "VLA is dead; World Action Models are the future." On October 3, Nvidia further expanded its open physical AI stack, releasing the Cosmos world model, Isaac Lab Arena, autonomous driving's Alpamayo, orchestration tool OSMO, and OpenUSD library together, with Caterpillar, LEM Surgical, and Neura Robotics as first users. On October 4, multiple humanoid robots from DEEP Robotics in Hangzhou, Zhejiang, took to the streets to test intersection traffic direction and tourist inquiries, and they were running even in the rain.
With such density, universities open-sourcing a full-stack foundation is equivalent to lowering the barrier to the direction by a notch, and making "learn the world from video, learn actions from trajectories" a more public route.
It is not meant for direct deployment
To be clear, OpenWAM is positioned as research infrastructure, not an out-of-the-box product. The official README recommends preparing about 640GB of training data, the environment size is about 6.5GB, and the repository is still actively changing. It suits teams that already have GPUs and data and want to research World Action Models on top of it, and is not suitable for small-scale deployment scenarios.
The follow-up points worth watching are also very concrete: how high the reproduction rate of this batch is, how many people are still submitting improvements on it six months later, and whether any team actually fine-tunes it into a product line. The launch-day buzz will fade; what remains is the code that is still being used.
Related articles
A Federal Judge Called Flock's License Plate Network “Indiscriminate Mass Surveillance”
One sighting of a car reveals little. A month of aggregated sightings from many agencies reveals a life.
Federal Prosecutors Say $300 Million in Nvidia Servers Went to China Through Malaysia
Enforcement cases measure the flow rather than stop it, and the transit route is the part policy has not solved.
Supabase Bought Turso Because Agents Need a Database Per Task
One database per agent only makes sense once the smallest unit of storage gets cheap enough to hand out like a session token.
Actors Got a Contract That Regulates Their Digital Replicas. The Hard Part Is Enforcement.
A contract can define what a digital replica is. Proving that one was made without consent is a different problem.