← Back to blog
AiAbout 6 min read

Reka Rho-1 Puts Understanding and Generation in the Same KV Cache

Published Oct 6, 2026
Reka Rho-1 Puts Understanding and Generation in the Same KV Cache

Most multimodal systems are pipelines. A language model plans, a diffusion model renders, a detector localizes, and a policy network drives the robot. Each handoff has a cost. State has to be serialized, moved between services, and re-encoded into whatever format the next component expects.

Reka AI's Rho-1 takes a different route. The 19-billion-parameter model, released on October 5 as a research preview, handles text, images, video, and robot actions inside one network, with everything represented as tokens in a single context window. There are no tool calls and no second model.

The demo makes the difference concrete. A user asks for a low aerial view of a red-and-white lighthouse on a rocky headland. Rho-1 renders it. The user asks for a box around the lighthouse; it draws one. The user asks to animate it as a drone fly-in; it generates video from the frame of the image it just made. The user asks for a snowstorm while keeping the camera move identical; it edits the video. Finally, the user asks what changed between the two clips, and the model answers in text.

Every step in that sequence depends on the model still having access to what it produced earlier. A conventional pipeline would have to pass the image and video between services and reconstruct the context at each boundary. Rho-1 keeps it in the same attention context, which is why the chain holds together.

Two streams, one cache

The architecture divides each transformer block between two expert weight streams that share an attention operation. An understanding stream processes language, internal reasoning tokens, and visual input, and a next-token head emits discrete output. A generation stream uses a flow-matching head to denoise continuous latent representations into images and video.

Both streams read from the same KV cache. Instructions, previously generated visual content, reasoning tokens, and motor commands all remain available. A special token signals when a response needs visual generation, which lets the generation stream continue from the accumulated state rather than starting fresh.

One clear glass sphere containing two separate glowing chambers, one warm amber and one cool blue

The model carries information in two formats on purpose. Discrete tokens handle text and symbolic reasoning. Continuous representations carry image latents, video frames, robot actions, and proprioception. Keeping physical signals continuous avoids the precision loss that comes from forcing joint positions and velocities into a discrete vocabulary.

That last detail is what makes the robot control claim more than a marketing line. The same weights that predict a camera image also drive robot movement, because both live in the same representational space.

Coming at generation from the other direction

Reka reports two variants. The base model uses 99 denoising steps and generates at a median rate of 0.79 times real time, with streaming output beginning after roughly six seconds. A distilled version called Rho-1 Flash uses eight steps and returns a 5.3-second clip in about one second. Flash edits re-render roughly one-second clips in about 1.1 seconds.

The streaming interface can extend a clip from its previous latent state and accept new instructions during generation. In one aerial demo, commands to bank left and bank right arrive half a second into the rollout. The resulting branches share the same opening frames before diverging. That is a real capability if it holds outside a controlled demo, because it means a director can steer a shot mid-flight instead of regenerating from scratch.

To work around the shortage of robot training data, Reka built an inverse dynamics model that extracts control signals from ordinary internet video. It trained the shared backbone on 320 H100 GPUs over about three months, which is modest compared with frontier runs.

The compute story deserves emphasis. Training a 19B model on 320 H100s for three months is a small run by frontier standards, where systems are trained on tens of thousands of accelerators. If Rho-1 delivers even a fraction of its architectural promise at that scale, it suggests that the next wave of multimodal capability will not require enormous capital, and that smaller labs can still make structural contributions rather than competing on compute alone.

Where this fits in the world-model race

Rho-1 arrives in a crowded week for models that try to represent how a scene evolves. InSpatio released InSpatio-World 1.5, which turns a single image, panorama, or video into a navigable 4D world and topped one dynamic-scene benchmark among real-time interactive methods. Nvidia open-sourced Lyra 2.0, which converts an image into an explorable 3D environment. Google and a set of Chinese labs are all working the same seam.

Reka's angle is different from the reconstruction crowd. Those projects build a scene you can move through. Rho-1 tries to build a model that can both describe a scene and change it on request, while also outputting the motor commands to act on it. If that works, the same checkpoint could drive a robot and narrate what the robot sees, without a separate policy network or a separate vision-language model.

The robotics claim is the most consequential and the least verified. Reka said the same weights that predict camera images also drive robot movement, and that it built an inverse dynamics model to pull control signals from ordinary internet video because robot data is scarce. If the approach holds, it points at a way around the biggest bottleneck in embodied AI, which is that real robot trajectories are expensive to collect and limited in variety. If it does not, the robot control feature is a demo clip.

The release also matters for how these systems are priced. Most multimodal products charge separately for planning, image generation, and video generation, because each is a different service. A single model that performs all three in one context window changes the unit economics for any product built on top. Whether that shows up as lower prices depends on how efficiently the unified model serves concurrent requests, which is exactly what independent testing would reveal.

The claims that are still open

Rho-1's performance figures come from Reka's preview and cannot yet be reproduced, because there are no public weights or API. Hardware configuration, batching, output settings, and compilation choices can change video latency by large factors, so the efficiency numbers need independent replication before they mean much.

The model has admitted limits. Native video is capped at 672x384. Long-horizon structural drift, video grounding, and edit stability are described as weak points by the company itself. Those are the same problems that plague every world model, and a 19B network is unlikely to have solved them outright.

The release fits a broader shift toward world models, where the goal is a system that can predict how a scene evolves and act on that prediction. Rho-1's contribution is architectural: it argues that understanding and generation should share weights and context rather than being stitched together. Whether that argument wins depends on what the next few research previews from other labs look like, and on whether Reka ships anything a third party can test.

Related articles