← Back to blog
AiAbout 6 min read

Satellite Photos Are Now Robot Training Data

Published Oct 7, 2026
Satellite Photos Are Now Robot Training Data

The fastest route to teaching a machine a new place may not be to drive it there. Two releases this month point at the same idea from opposite directions. One turns satellite imagery into simulated worksites so robots can practice before arrival. The other fixes the geometry underneath video world models so that what a robot imagines matches where things actually are.

Bonsai Robotics, which builds autonomy software for farm and mining equipment, introduced Bonsai World on October 2. The system starts with satellite imagery of a farm, mine site or other rough terrain and turns that flat view into a structured 3D simulation. Machines can then rehearse real paths inside it. The company says it can layer in conditions the fleet has not physically met, including dust, debris, animals, vehicles and shifting ground.

The compute behind it is a cross-vendor stack. Google's Gemini vision model reads the satellite image and builds a structured map. Bonsai's own world model, post-trained on more than 50 million real-world samples collected across over a million acres, generates the ground-level views. Training runs on Google A2 virtual machines with Nvidia A100 GPUs, while on-demand environment generation uses Google G4 machines with RTX PRO 6000 Blackwell server GPUs. Nvidia's Cosmos models sit underneath.

The commercial argument is about bring-up time. Getting autonomy software working on a new crop, a new machine or a new site normally means collecting field data and iterating on location, which is slow and expensive. If the system can rehearse against a plausible reconstruction first, the machine arrives closer to ready. Bonsai says the approach reduces field collection rather than replacing it, which is the honest version of the claim.

The geometry problem nobody sees

Simulation from imagery only works if the 3D it produces is truthful. That is where a paper published on arXiv this month gets interesting. A team from Czech Technical University and Inria released DepthWorld, accepted at the Conference on Robot Learning, and it starts with an admission that undercuts a lot of demos: current video world models are trained on RGB alone, and they produce rollouts that look right frame by frame without composing into a consistent 3D world.

The failure is easy to picture. Give an RGB-only world model a scene with an open wardrobe and it cannot tell the open door from a solid surface, so it simulates a robot arm colliding with empty space. If you are using the model to evaluate a policy, that kind of error produces a signal that is worse than useless, because it looks like a real failure and is not.

The team's first fix is data. They built a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling every episode collected from the same physical robot so the model can recover that robot's shared kinematics alongside each scene's camera extrinsics. Applied to DROID, the largest open teleoperation dataset, this produces DROID-3D: dense metric depth and recalibrated multi-view extrinsics across more than 70,000 episodes, with reprojection error under 0.7 pixels on 90 percent of episodes for external cameras.

The second fix is architectural. Rather than re-initializing a pretrained video model to add depth, which risks destroying the visual and motion priors it already learned, they tile RGB and depth side by side in a wider latent grid and extend the position embeddings to cover it. The pretrained component stays untouched. The reward is a result that surprised me when I read it: adding depth as a second prediction target improved the RGB prediction itself, by 1.48 dB PSNR at equal training budget.

Why that number matters beyond the paper

A world model is only useful for planning if its geometry holds up over a long rollout, and that is where the comparison with Nvidia's PointWorld got interesting. PointWorld was more accurate on moving objects at a short horizon, but DepthWorld degraded much less over time, going from 8 to 9 millimeters of whole-scene error where the other went from 8 to 18. For a system that plans minutes ahead, the curve matters more than the starting point.

Put Bonsai World and DepthWorld together and a pattern shows up. The bottleneck in physical AI training stopped being compute or even model capability. It became the quality of the synthetic world, and specifically whether the geometry inside it is metric, calibrated and stable. That is a data engineering problem before it is a modeling problem. The DROID-3D contribution is a pipeline rather than an architecture: it recovers millimeter-accurate camera poses from a dataset that was never designed to have them, and it does so by pooling episodes instead of trusting any single one.

What improves, and what does not

The gains are real and narrow. Bonsai World works on structured outdoor environments where satellite imagery is available and terrain is the dominant variable. That covers agriculture and surface mining well. It covers a cluttered kitchen or a hospital corridor poorly, because those spaces are not visible from orbit and their difficulty comes from objects, not ground. The company's own framing, rugged and unstructured environments, is a scope statement as much as a market description.

DepthWorld has the opposite limitation. It is strongest where you have many episodes from one robot and can pool their calibration, which is exactly the setup in a research lab and less common in a deployed fleet with mixed hardware. The paper's own benchmark is a single open dataset.

There is also a verification gap worth naming. Bonsai did not publish independent field performance measurements or open model weights with its announcement. The world model is a claim until someone tests it outside the company. DepthWorld's code and dataset are the opposite case: a published paper, an accepted venue and a reproducible pipeline, which is why it will probably influence how other teams build their own world models even if nobody uses this specific one.

The part that should worry robotics teams

If synthetic environments become the default way to pre-train autonomy, then the quality of the simulation becomes a single point of failure across many deployments. A bias baked into the generator, whether about lighting, terrain or the distribution of obstacles, propagates into every machine trained on it, and it does so invisibly, because the machines that fail never make it to the field to reveal it.

That is the argument for keeping some real-world collection in the loop even when the synthetic version looks convincing. Field data earns its keep through the examples it provides and, more importantly, through the check it puts on the simulator. Both of this month's releases move the field forward on that front, one by making simulated environments cheap to generate and the other by making the geometry inside them honest. Neither removes the need to eventually drive the machine out and see what it does.

Related articles