← Back to blog
AiAbout 6 min read

One Photo, a Walkable World: InSpatio-World 1.5 Turns Images Into Spaces

Published Oct 3, 2026
One Photo, a Walkable World: InSpatio-World 1.5 Turns Images Into Spaces

Most image and video models give you a fixed frame or a fixed clip. You look at what the model decided to show you. A Chinese company called InSpatio released a model this month that takes a single photo, or a panorama, or an existing video, and turns it into a space you can move around in.

InSpatio-World 1.5 went live at the Shanghai International Audiovisual Arts Carnival in late September, and the company open-sourced the code under Apache 2.0. The demo video is the clearest explanation of what is different. It starts inside the covered walkway of a Chinese palace. The camera pushes forward, passes a railing and a set of steps, and turns into a courtyard. Pavilions that were hidden behind columns come into view as the viewpoint shifts. What the model produced is a space that keeps rendering as you keep exploring it, rather than a clip with a chosen camera path.

A misty corridor of stone columns receding toward a doorway of bright light

Three things changed in this version

The first is what the model accepts as input. Earlier world models mostly wanted a single image. Version 1.5 takes a single image, a set of images, a panorama, or a video. That matters because it sets the entry point. A single photo is a consumer-grade starting point. Multi-image sets and existing video are what a film or advertising team already has in hand, and accepting them widens who can use the tool.

The second is real-time exploration. Once inside the generated space, you can keep moving the viewpoint, go around occlusions, and reach areas that were not in the initial frame. Every time the viewpoint changes, the model has to fill in the parts of the scene that were previously invisible, and do it consistently with what you have already seen. That is a harder problem than generating a pleasant clip, because the model is being asked to maintain a mental model of a place rather than a sequence of pictures.

The third is spatiotemporal consistency, which is the dividing line between a world model and a video generator. If the camera leaves a room and comes back, the room should not have rearranged itself. InSpatio says it strengthened this in 1.5 to support longer exploration paths and more complex scenes, and the demo leans on that claim by repeatedly revisiting areas from new angles.

The numbers behind the demo

InSpatio describes the model as compact, at 1.3 billion parameters, and reports a score of 68.72 on WorldScore-Dynamic, which it says ranks first among evaluated real-time interactive methods, at up to 24 frames per second. Those are the company's own figures, and independent reproduction is still thin, so treat them as a starting position rather than a settled result.

The technical approach has a couple of moving parts worth naming. For video inputs, the pipeline uses Depth Anything 3 to estimate missing depth and per-frame camera intrinsics and extrinsics, which is what lets it reconstruct a scene from footage that was not shot as a 3D capture. A spatiotemporal autoregressive process is what maintains a persistent world state as the view changes, rather than regenerating the scene from scratch each time.

Where a walkable scene earns its keep

The most immediate use is the one the demo points at. In film and advertising, a camera position is usually locked the moment you shoot. If you want another angle, you reshoot, and reshoots are expensive. A world model lets a team keep looking for camera positions after the fact, pulling coverage out of footage that already exists. InSpatio calls out advertising bullet-time as a direct example: instead of synchronizing a rig of cameras around a subject, you feed in a set of images and design a new camera path around the same subject.

The second use is training data for robots. One problem in embodied AI is that most of the video the internet provides is shot in the third person, while a robot sees the world in the first person. InSpatio says the model can take third-person footage of an operation and predict a first-person view closer to what a robot would perceive, and can re-observe the same process from different directions with the occlusion relationships changed. If that works at scale, it is a way to turn available video into training material for robots without shooting new footage for every environment.

Beyond those, the company is pointing at tourism and cultural experiences, digital twins and industrial simulation, the places where a space that behaves like a space is more useful than a picture that looks like one.

A crowded race, with different definitions of winning

InSpatio is not alone in chasing this, and the competing approaches disagree about what a world model is even for. Some labs optimize for cinematic output, generating a beautifully consistent minute of footage with a single narrative camera move. Others chase interactivity, prioritizing the ability to move freely and have the scene respond without a fixed path. A third camp, closer to gaming, wants real-time engines that hold a world in memory as a navigable place.

Those goals pull in different directions. A world model tuned for cinematic quality can afford to spend compute on a predetermined sequence and make every frame look right. One tuned for interaction has to render whatever the user looks at next, which pushes toward speed and toward a memory of the scene that survives arbitrary movement. InSpatio-World 1.5 is clearly in the interactive camp, and its choice of a small 1.3-billion-parameter model is a bet that staying real-time matters more than winning a still-frame beauty contest.

The comparison that matters for buyers is against the other navigable systems, and here the field is early enough that most claims are self-reported. A benchmark like WorldScore-Dynamic tries to put a number on the slippery part, whether the world stays consistent as you move through it, but a benchmark only measures what its designers chose to measure. Consistency across a long exploration, fidelity to the original photo, and frame rate all pull against one another, and every system will pick different tradeoffs.

That is the strongest reason to watch the open release. When dozens of researchers stress-test the model on their own inputs, the picture of where it holds and where it breaks will be far more useful than the launch demo, and it will let the rest of the field compare approaches on shared ground instead of on competing highlight reels.

Why open-sourcing is the strategy, not the giveaway

InSpatio released the code alongside the model, and is teasing a new Topos-Pro 1.0 for invited testing. The reasoning is the one that shows up across spatial intelligence: the useful applications are scattered across film, tourism, robotics and digital twins, and no single company can cover all of them. Publishing the base capability lets researchers verify the limits and lets partners build the vertical layer on top. It is a way to be the ground floor rather than the whole building.

The risk is the usual one with an early open release. The demo is impressive and the benchmark is self-reported, and the gap between a controlled demo and a tool that holds up on a messy real-world input is often wide. The next few months, as independent researchers and creators get their hands on the code, will show which parts of the promise survive contact with other people's data.

Related articles