← Back to blog
AiAbout 7 min read

Unitree and ZDTaichu Race to Be the Robot's Spatial Brain

Published Oct 2, 2026
Unitree and ZDTaichu Race to Be the Robot's Spatial Brain

Robots have been good at bodies for a while. A humanoid can walk, balance, and grip. What it has been bad at is the part that decides what to do next, especially when the room is not the room it was trained in. Move a cup, change the workbench, hand the robot a task halfway through, and the demo falls apart.

Two Chinese releases in September went after that part. Unitree showed off UnifoLM-WLA-1.0, a six billion parameter humanoid foundation model trained on roughly 2,500 hours of real robot data, claiming seven open-source state-of-the-art results on embodied reasoning. ZDTaichu, spun out of the Chinese Academy of Sciences and the Wuhan AI Research Institute, open-sourced ZDTaichu5.0-9B on 15 September, a nine billion parameter multimodal model that took first place in eight of nine international spatial understanding benchmarks within its parameter class.

Neither is a robot. Both are attempts to build the brain one runs on, and both were released rather than shown behind a demo wall.

Unitree's bet on generality

The number that stands out in Unitree's release is 64. A single model claims to coordinate 64 different tasks, from folding towels and putting dishes away to making a bed, taking out the trash, and loading a washing machine. The point of that list is not any individual task. It is that the same weights handle all of them.

That has been the central problem in robotics for years. Most demonstrations are single-task, single-scene, and heavily tuned. They look impressive and they do not transfer. A policy that folds a towel in one lighting condition does not fold a differently coloured towel on a different table. Scaling a robot workforce on that model means teaching it every variation, which is not scaling.

Unitree's approach leans on an embodied reasoning base called UnifoLM-ER-1-4B, built on top of Qwen3-VL-4B and trained on more than five million embodied reasoning samples. On 16 multimodal perception and understanding benchmarks, it leads the field in seven of them. On two spatial tests, Where2Place and EmbSpatial, it scores 82.0 and 88.9 against GPT-6 Astra's 69.0 and 83.3. Those are specific comparisons against a closed model, which makes them easier to argue about than a general claim of superiority.

Unitree says it will open the model, the code, and the dataset. That last item is the one that matters. A robot policy with open weights but a closed dataset cannot be reproduced, only used. Releasing the data is what lets another lab check the claim or build on it.

ZDTaichu's bet on spatial reasoning

ZDTaichu5.0-9B attacks the problem from the perception side. The model handles text, single images, multiple images, long video, and arbitrary resolutions, and its selling point is spatial reasoning: understanding where objects are, how they relate, and how a scene changes when the viewpoint moves.

The benchmark gap it highlights is on MindCube-tiny, where it scores 78.27, ahead of Qwen3.5-9B by more than 20 points and STEP3-VL-10B by more than 15. In one demonstration, asked to find the second silver box from the left, it outputs the correct coordinates while other open models of similar size all point to the wrong place. In a cross-viewpoint test, only this model keeps the reference frame straight when the camera angle changes.

Those are the failures that break real robots. A model that recognises a cup has answered what the object is. A model that knows which way the handle faces, and whether there is room to set it down beside it, has started to answer what can be done with it. When a task runs long, those judgments have to chain: object picked up, identity remembered, container opened, state updated. One missed step and the rest of the task drifts.

ZDTaichu's engineering trick is adaptive loop reasoning. Easy questions get less computation. Hard ones, like spatial transformations and object relations, get multiple internal reasoning passes without calling out to an external tool, which keeps the token cost from ballooning on the simple majority of inputs.

The part of the release that may matter most is what it does not gate. Alongside the weights, ZDTaichu published its entire spatial multimodal data production pipeline, covering pretraining, supervised fine-tuning, and reinforcement learning. Institutions and companies can reuse it to turn their own factory-floor or lab data into training samples. That answers a complaint that has followed open model releases for two years: you get the weights, but you cannot adapt them to your own data because the recipe stays private. The model has been running at embodied training grounds in Beijing, Foshan, and Qingdao.

Two routes to the same gap

Put the two releases side by side and the difference is instructive. Unitree works from the action side outward: take a policy that has to perform tasks, train it on 2,500 hours of real robot motion, and measure whether it generalises across 64 jobs. ZDTaichu works from the perception side inward: take a multimodal model that has to understand a scene, train it on spatial benchmarks, and measure whether it keeps the geometry straight when the viewpoint moves.

A robot needs both. It has to know where the cup is, and it has to know how to pick it up. The two labs are attacking the same gap from opposite ends, and the fact that both published the same month suggests the field has agreed on what the gap is: the middle layer where perception turns into action, and where a task that ran for eight seconds turns into one that runs for eight minutes.

The data question separates them again. Unitree trained on real robot motion, which is expensive to collect and rare. ZDTaichu leaned on its data pipeline and synthetic spatial tasks, which scale differently and carry a different risk, that a model gets good at test scenes and stays there. Whichever route produces a policy that survives contact with a real warehouse will settle which data source was the right bet.

The industry context

Both releases land in a field that has reorganised its assumptions recently. Google's Gemini Robotics 2 pitches "one brain for any robot." Nvidia's Jim Fan has argued that vision-language-action models are dead and world action models are the successor. Gartner's September report on China's AI trends put physical AI and world models among the directions that have become structural requirements rather than experiments.

That reframing is why a spatial benchmark matters more than a chat score. The competition in embodied AI has moved from how well a model talks about the physical world to whether it can act in one, and the metrics have followed. Coordinate accuracy, viewpoint consistency, and task completion across scenes are the new scoreboard.

What to watch

The honest caveat on both releases is that benchmark leadership in spatial reasoning is not the same as a robot that works in a warehouse. A model can place objects correctly in a test set and still fail on a real arm with a stuck gripper or bad lighting. The gap between reasoning and acting is where most robotics promises have died.

What makes these two worth tracking is the combination of open weights and open data pipelines. If a lab outside Unitree or ZDTaichu can take either model, retrain it on its own hardware, and publish the result, the claims get tested in public. If the releases stay benchmarks and demos while competitors keep their data closed, the field will stay where it has been: a lot of impressive video and very few robots doing useful work on a Tuesday.

Related articles