Fei-Fei Li's Atlas Turns a Photo Into a Room You Can Walk Through

World Labs put Atlas into early access on 1 September. The company, co-founded by Fei-Fei Li, calls it an omni world model for spatial intelligence, and the demonstration is a few ordinary phone photos of a room going in, and an interactive 3D scene coming out. You can move a virtual camera anywhere in it. You can read depth from it. You can hand the geometry to a robot and use it as a place to practise.
Li has been arguing for years that the next frontier of AI is not language but the physical world. Her term is spatial intelligence, and the claim is that a model trained only on text and flat images is missing something fundamental. Language models describe a world they cannot see. Atlas is the clearest attempt yet to build one that can.
What Atlas does
World Labs describes Atlas as a multimodal autoregressive diffusion transformer, a single architecture that takes in text, images, video, and 3D data inside one spatial framework and generates across the same set. The outputs are broader than a video generator's: camera-controlled images and video up to 1440p for clips of about a minute, plus novel viewpoints, depth maps, point clouds, and 3D Gaussian splats.
The volumetric output is what separates this from a text-to-video model. A video generator predicts the next frame. It does not necessarily know where anything sits in space or how a scene looks from an angle the camera never occupied. Atlas carries an internal 3D representation, so an environment holds together as a viewer moves through it. The difference between a plausible clip and a navigable space is the whole point of calling this a world model.
From a sparse input, in some demonstrations just a few frames of phone footage, the model reconstructs a scene well enough to place a virtual camera inside it. World Labs reports a sparse-view reconstruction error of 25.3, below the best specialist model at 28.7. In blind evaluations it commissioned, human raters preferred Atlas's adherence to a requested camera move over MiniMax H3 75 percent of the time, Gemini Omni Flash 81 percent, and Seedance 2.5 94 percent.
Those are the company's own numbers from evaluations it ran, which is worth saying plainly. Atlas is in early access, so nobody outside the partner group can reproduce them yet.
Real-to-sim, and why robots care
The obvious commercial applications are visual effects, games, and virtual production. A filmmaker can stage a shot after the shoot. The framing World Labs leans on, though, is robotics.
The workflow is called real-to-sim. You capture video of a real space, and Atlas converts it into a 3D simulation where a robot can be trained or tested without the cost and risk of operating the real thing. A robotics team that wants a thousand variations of a warehouse aisle can scan the aisle once and generate the variations instead of filming each one.
That is a narrow but genuine pain point. Training data for robots is expensive because collecting it means operating robots. Simulation is cheap but usually requires someone to build the environment by hand, which is slow and often looks nothing like the real place. A model that turns a real room into a simulation file collapses two expensive steps into one. It also explains why Nvidia and AMD are investors. Both sell the compute that training and simulation run on.
The disclosure problem
There is a gap between the ambition and the paper trail. World Labs published no research paper, no model card, no parameter count, and no training compute figure alongside Atlas. It disclosed no price and no general availability date. For a system making comparative claims against well-documented competitors, that omission is unusual, and it makes the blind-evaluation results impossible to check.
Skeptics have raised a sharper question: whether Atlas genuinely simulates the world or renders convincing views of it. Those are different capabilities, and the distinction matters for every use case that depends on physics. A model that renders a room from a new angle is useful for a film. A model a robot learns to walk through needs the objects in it to behave the way objects do, or the policy learns something that only works in the model.
World Labs says Atlas will power future versions of Marble, its existing product. That gives it a commercial home, which is more than most research announcements have. Whether the geometry holds up beyond curated demos is the thing early access partners will learn first.
The Chinese counterpart
Atlas is not the only world model with a city-sized ambition. On 10 September, the Chinese mapping company Amap released ABot-Earth 0.7, which it calls the world's first 3D-native city world model. It takes satellite imagery or a text description and turns it into a walkable, interactive 3D scene, generating at multiple scales from a planet down to a street-level landmark inside one model. Amap says it can produce a kilometre-scale 3D city scene in about ten minutes on a consumer GPU, in the 3D Gaussian splatting format, and that the capability already runs in its Flight Street View product.
The two point at the same idea from opposite ends. Atlas works from a handful of photos upward, reconstructing a room with high fidelity. ABot-Earth works from satellite data downward, generating a city at scale. Together they sketch the range a spatial model can cover, and they also show how differently the two markets ship: one through a partner programme with no public benchmarks, the other embedded in a consumer product where anyone can look at the output.
What spatial models still have to prove
The field has momentum. Meta's V-JEPA 2 was trained on more than a million hours of video. Google's Genie 2 generates interactive 3D environments, and Gemini Robotics works on reasoning and manipulation in physical space. Li's argument, that geometry and perspective have to be native to the model rather than inferred from text, has moved from a research position to a product category.
What none of it has settled is whether a model that produces coherent geometry is the same thing as a model that understands a physical space. Reconstruction is measurable. Simulation is harder, and it is the one robotics actually needs. Atlas makes a strong case for the first. The next year of partner results will decide how much of the second comes with it.
For designers and developers, the near-term effect is more modest than the launch language suggests. A tool that reconstructs a room from three photos and lets you fly a camera through it is genuinely useful, and it will show up inside creative software before it shows up inside a robot. That is how spatial models reach most people first, through the work of making a scene rather than the work of navigating one.
Related articles
A Wheeled Semi-Humanoid Finished an Hour of Laundry Without Help
Individual tasks can succeed while a workflow still fails. Dyna changed the metric.
LTX 2.5 Wants to Render Your Blocky Blender Draft Into a Finished Shot
You do not control what happens in text-to-video. This tries to fix that.
ServiceNow Turns Agent Failures Into Training Data
Generation without verification is noise. The gates are the product.
One Framework for Language and Vision: Horizon's 1.6B Open Model
A bet that the bridges between language and vision were never needed.