← Back to blog
AiAbout 6 min read

Iris-3B Generates Pixels Without a VAE

Published Oct 11, 2026
Iris-3B Generates Pixels Without a VAE

Almost every popular image generator works the same way under the hood. It does not draw pixels directly. It first compresses an image into a smaller, abstract representation using a component called a variational autoencoder, or VAE, generates in that compressed space, and then decodes the result back into pixels. This two-step design is efficient, and it is also where a lot of the small flaws come from. Fine text smears. Thin lines wobble. Details get lost in the round trip.

In October 2026, a small lab called Speridlabs released Iris-3B, a three-billion-parameter model that skips the round trip entirely. It generates pixels directly. The weights are open under an Apache 2.0 license, and the whole thing is built on a flow-matching transformer. It is not the largest or the flashiest image model of the month. It is arguably the most interesting one for what it says about where image generation might be heading.

Why skipping the VAE matters

The VAE exists to make generation cheap. Working in a compressed latent space cuts the compute needed by a large factor, which is why nearly every diffusion model uses one. The trade-off is that the encoder throws information away before generation even starts, and the decoder has to guess it back. At low resolution the loss is invisible. At high resolution, or on images with dense text and precise edges, it shows up as the familiar mush that generators produce around small details.

Generating in pixel space removes that bottleneck because there is no compression step to lose anything. The cost is that you are now doing the work on a much larger grid of values, which is expensive. Iris-3B's claim is that it is small enough to make that expense manageable. Whether it holds up in practice is something the community will test, but the direction is clear: as efficiency improves, the reason to accept the VAE's quality penalty gets weaker.

A translucent cube of pure light assembling itself from a cloud of glowing dots

The second trick: a generative prior for vision tasks

The more unusual part of Iris-3B is what else it does. The model applies its generative prior to detail-critical vision tasks, including depth estimation and image restoration. In plain terms, a model trained to produce images can also be asked to recover information from them: estimate how far away things are, or clean up a degraded photo.

That is a meaningful combination, because it points at a broader trend. For a while, generation and understanding were treated as separate problems with separate models. Lately they have been merging. A system that can produce a plausible image already understands a great deal about how scenes are structured, how light falls, and how objects relate in space. Reusing that understanding for analysis tasks is cheaper than training a dedicated model for each one.

What a small open model is actually for

Iris-3B will not top any leaderboard against the giants. Three billion parameters is a fraction of what the commercial leaders run, and a general-purpose model at that size will lose on the hardest prompts. Judging it by that standard misses the point.

Open models at this scale earn their place in three ways. They run on hardware people already own, so there is no per-image bill and no data leaving the building. They can be fine-tuned for a narrow job, which is often where small models beat large ones: a model trained specifically on a company's product photography will outperform a general model on that exact task. And they are inspectable, which matters for teams that need to explain where an output came from.

The Apache 2.0 license is the detail that makes all three possible. A permissive license means a startup can build a product on top of Iris-3B without negotiating terms or worrying about a future price change. For a research lab trying to get its architecture adopted, that is the fastest path to relevance. The model does not need to win the benchmark. It needs to be the thing other people build with.

What flow matching changes, in plain terms

Iris-3B is built on a flow-matching transformer, and the idea is easier to grasp than the name suggests. Older image models learn by reversing a process that adds random noise to a picture, step by step, until the image disappears. Generation then runs that process backward, denoising step by step until a picture emerges. Flow matching takes a more direct route. It learns a straight path from noise to image rather than a wandering one, which in theory needs fewer steps and less compute to reach a good result.

That matters for a pixel-space model specifically. Working on raw pixels is expensive, and the usual way to control that cost is to make each step efficient. A straighter path means fewer steps, which is part of how a three-billion-parameter model can attempt a job that larger systems handle with the VAE shortcut. The architecture choice and the size of the model work together. Neither one alone would be enough.

Independent testing will decide whether the trade-off pays off. If pixel-space generation matches latent generation at a similar cost, the VAE stops being a default and becomes a choice. If it does not, Iris-3B is still a useful data point about where the wall sits.

One caveat applies to every open model release, and Iris-3B is no exception. A permissive license covers the weights, not the training data or the outputs. Teams that plan to ship commercially still have to think about what the model learned from and what rights attach to a generated image. That is a legal question the license does not answer, and it is the same question every open model raises. The freedom to run the model yourself is real and valuable. It is not the same as freedom from liability.

The bigger picture

The image field in October 2026 looks crowded from the top. Nano Banana 2.1, GPT Image 2.5, FLUX 3, and Seedream 5.0 all shipped recent updates, and the gains at the frontier are measured in small margins. Underneath, the more interesting movement is architectural. Pixel-space generation questions an assumption that has held since the first diffusion models: that you must compress before you create. And open, small models keep proving that the fastest way to spread an idea through the field is to give it away.

Iris-3B may end up a footnote, or it may be the version of an idea that others copy. Either way, the two questions it raises are the ones worth watching. Can you generate images at full detail without a lossy intermediate step, and can one model both create and analyze? If the answers turn out to be yes, the next generation of image tools will be built on a different foundation than the last one.

Related articles