Unitree Open-Sources Its 6B Humanoid Foundation Model: One Model Runs Both Hands and Feet, and Embodied AI Starts Keeping Score by the Job

On October 3, Unitree released its next-generation general-purpose humanoid robot foundation model, UnifoLM-WLA-1.0, on GitHub and Hugging Face. It has 6 billion parameters and was fine-tuned on roughly 2,500 hours of real robot data. News of it didn't take long to spread through China's robotics circles, but the detail most easily overlooked is that it stuffs both "using hands" and "walking" into a single model.
Over the past year, the highlight moments for humanoid robots have mostly stayed inside demo clips: backflips, dancing, carrying a glass of water. But to actually get into factories, stores, and homes, the test changes. Whether a single motion is impressive enough is one question; whether one model can handle both fine tabletop manipulation and whole-body locomotion tasks is another. The number Unitree gives this time is 64 tasks — 10 whole-body tasks and 54 tabletop tasks — and it supports parallel grippers plus dexterous hands in two different configurations.
The tech stack is worth a look too. It starts with an embodied reasoner built on Qwen3-VL, layers on optical-flow and VQ-VAE-based future dynamic region prediction and residual VQ action discretization, and finishes with an MMDiT action expert. In plainer terms: first let the model understand what's about to happen in the frame, then break the "action to be taken" into reusable building blocks, and finally hand those to the module responsible for execution to assemble.

Why "One Model for Hands and Feet" Is a Watershed Moment
In the engineering practice of embodied AI, hands and feet have long been two separate teams. Grasping, inserting, and screwing are high-frequency, small-amplitude motions that demand precision and force control; walking, turning, and going up and down slopes are low-frequency, large-amplitude motions that have to handle balance and terrain. Splitting the two across different models has the advantage that each path can be tuned to the limit, at the cost of state handoffs whenever you switch — and every handoff can lose information and add latency.
By compressing 64 tasks into a single 6B model, Unitree is essentially betting that the payoff from a unified representation outweighs the payoff from optimizing each path separately. That bet isn't cheap on general-purpose GPUs, but it makes sense on a robot: in real-world deployment, VRAM, compute, and power all have hard constraints, and one fewer model means one less cost.
Unitree Isn't the Only One Pushing This
Line up a few things from around October and it becomes clear this isn't an isolated release.
BAAI open-sourced RoboBrain-X0 on September 29, extending large-model capabilities into robot perception and motion control. Open-source embodied AI projects had previously mostly stopped at the level of simulation environments and datasets; the emergence of a model aimed directly at action policies shows this line of work is moving from papers to reproducible engineering assets.
On October 4, several humanoid robots from Hangzhou-based DEEP Robotics took to the streets to test intersection traffic direction, order maintenance, and tourist inquiries. The test site was a real intersection rather than a demo booth, and they ran in the rain too. The value of such tests isn't how smooth the motions are, but that they put a humanoid into a real environment that won't make way for it.
Go back a bit further, and seven universities — the National University of Singapore, Tsinghua University, Peking University, and others — jointly open-sourced a world-action model foundation called OpenWAM. It tackles that old rift between simulation and real robots: traditional simulators tend to compute physics and vision separately, whereas OpenWAM has the model learn directly from data "how actions change the world," then use that to predict object positions, scene semantics, and task states. The official README recommends roughly 640GB of training data; it's positioned as research infrastructure, not an out-of-the-box product.
From "Can It Move?" to "Can It Hold a Job?"
The embodied AI narrative is switching tracks. A couple of years ago the contest was whether a body could walk and whether its moves were cool enough; now it's whether a robot can work for hours on end without a hitch.
This shift has a very concrete metric: the average time between two human interventions. One company has publicly walked through a piece of "reliability arithmetic": a single laundry cycle strings together roughly 79 subtasks, and at a 95% success rate per step, the odds of completing the whole round with no intervention are only about 1.7%. A 95% success rate per subtask looks high on its own, but multiply them together and it collapses. That's why the industry is starting to change its optimization target from per-task success rate to how long a complete workflow can hold up.
That's also where the value of Unitree's open-sourced 6B foundation model lies. Its significance is in providing a reproducible common starting point. Developers can later fine-tune on this base, swap hardware, and add tasks without having to accumulate data from scratch. Competition in embodied AI is shifting from "build a body that can move" to "can the brain be reused across hardware and tasks."
Final Thoughts
Open-sourcing a 6-billion-parameter humanoid foundation model shows no short-term commercial return, and the data cost is real. But it solves something more fundamental: giving researchers and developers a common reference to calibrate against.
For robots, the real test comes in those thousands of hours in factories and stores; the launch event is just the opening act. Whoever turns "continuously usable" into something real is the one who has truly entered the game.
The things to watch are very concrete: how high the reproduction rate is, how fast iteration moves in real projects, and how many people are still submitting improvements six months from now. The buzz on launch day will fade; all that's left is the code still in use.
Related articles
The Eighth Circuit Paused Minnesota's AI Nudification Ban. The Real Fight Is Over Who Is Liable.
The ruling is not a decision that the statute is unconstitutional. It pauses enforcement pending appeal.
California Signed 13 AI Bills in One Day. SB 1000 Deleted the Million-User Threshold.
For a category of startups that assumed they were too small to be regulated, that assumption just expired.
Three Companies Built Video Products on MiniMax H3 in One Day. The License Excludes the US.
A product sold to US enterprises by a US company, built on weights whose license excludes the US, is a question a statement from MiniMax would settle.
Claude Sonnet 5.5 Is 30 Percent Cheaper. That Is a Procurement Story, Not a Model Story.
Vendor discount claims are usually measured against a representative workload, and yours is not representative.