← Back to blog
AiAbout 6 min read

PixelUMM and Sana Point at the Same Idea: Take the Scaffolding Out of Vision Models

Published Oct 7, 2026
PixelUMM and Sana Point at the Same Idea: Take the Scaffolding Out of Vision Models

Two NVIDIA research releases landed within days of each other, and they share a thesis. PixelUMM appeared without fanfare on Hugging Face on October 1 at 21:40 UTC, a 15.2-billion-parameter checkpoint with no blog post, no press release, no keynote slide. Sana, from NVIDIA, MIT and Tsinghua researchers, has been public longer and takes the opposite approach to the same plumbing problem.

PixelUMM: a unified model with no encoder

The repository name says what the project is: nv-tlabs/PixelUMM, described in its own one-line summary as "encoder-free unified image and video understanding and generation." It sits on a Qwen3-8B backbone with an Apache-2.0 repository license.

The "encoder-free" part is the interesting claim. Most visual AI systems today chain together separate components: a vision encoder to turn an image into tokens a language model can read, a VAE to compress pixel data into a latent space a diffusion model can work in, and a visual transformer to handle the sequence. PixelUMM was built to operate without those intermediary stages, which is why its own project page still reads "preview" while the weights are downloadable.

That gap between an artifact that exists and a vendor that has said nothing is the whole story of the release. There is a GitHub repo, an arXiv preprint numbered 2609.38597 dated September 29 with authors from NVIDIA and the University of Waterloo, and a checkpoint split across 128 files with a hidden index that loaders refuse to run without. Whether NVIDIA considers PixelUMM announced is a question the company has not answered.

The release pattern matters for anyone tracking this field. Research that once arrived with a blog post, a demo page and a coordinated press cycle now sometimes arrives as a repository and a paper. The artifact is public and citable while the vendor's own communication is silent, which makes it possible to use the model and impossible to cite an official benchmark for it. Teams evaluating it are working from a paper and their own tests.

Sana: compress the cost chain instead of one module

Sana attacks the same layer of the stack from the other direction. High-resolution text-to-image gets expensive for a reason that has little to do with pixel count: the number of latent tokens entering the diffusion transformer grows quickly with resolution. Standard self-attention has to relate every token to every other token, so cost, memory and latency all rise together.

On NVIDIA's own 1024-pixel comparison, FLUX-dev carries 12 billion parameters, runs at 0.04 samples per second and takes 23 seconds per image. Pushing toward 2K or 4K, shrinking the model or cutting sampling steps does not fix the underlying problem.

Sana compresses the whole chain rather than one component. A 32x deep compression autoencoder cuts the number of latent tokens. Linear attention reduces the cost of each transformer layer. An efficient solver and few-step distillation cut the number of sampling passes. Slicing, offloading and low-bit quantization reduce deployment memory. The published models include 0.6B and 1.6B versions aimed at up to 4K, with Sana-1.5 extending to 4.8B. In the official 1024-pixel comparison, the 0.6B version runs at 0.9 seconds of latency.

The chain-compression approach has an attractive property that a single-module fix does not. Because every stage contributes, a team can apply the parts that fit its hardware and skip the rest. Someone with a workstation and no quantization experience can take the efficient solver. Someone deploying at scale can stack slicing, offloading and low-bit quantization to fit a much smaller card. The trade-offs are documented per stage rather than bundled into one all-or-nothing decision.

Sana's lineage also shows how quickly a research result becomes a product surface. The original model targeted 4K text-to-video output. The repository has since grown to include Sana-Sprint, video generation, ControlNet, LoRA, quantization, ComfyUI integration and an online service. That is the same arc the image tools followed: from a paper, to a repository, to a node in the interface people already use.

What the two approaches share

Put the releases side by side and the direction of travel is clear. Both are trying to remove the intermediary machinery that has sat between pixels and models since the field's early days. PixelUMM asks whether the encoder is necessary at all. Sana asks how much of the compute chain can be compressed before quality breaks.

The trade-off is the same in both cases, and neither lab is hiding it. Removing an encoder means the model has to learn what the encoder used to provide, which costs training compute and can cost quality on tasks the encoder handled well. Compressing the chain aggressively means the quality ceiling moves, and Sana's own materials are careful about the hardware and measurement conditions behind those latency numbers.

The reason both labs attack this layer is that the intermediaries have become the expensive part. A vision encoder and a VAE are trained separately, tuned separately and served separately. Each one is a component that can drift out of sync with the rest of the pipeline, and each one adds latency at the front of every request. Removing them is an architectural simplification rather than a research trick, and it pays off in every deployment.

Why it matters for anyone building with images

For practitioners, the practical consequence is lower hardware floors. NVIDIA's Sana work is part of a broader pattern this year: models that used to need a data-center card are being reworked to fit on consumer hardware, and the open community has started treating a 6GB VRAM floor as a requirement rather than a bonus.

There is a second consequence that shows up in architecture diagrams. When the encoder and the VAE stop being separate services, a pipeline that used to have three models to version, monitor and pay for becomes one. That is less visible than a latency number, but it is what changes the maintenance burden on a real product.

For teams building on top of these models, the practical question is which simplification they can adopt first. The encoder-free approach promises a cleaner architecture but requires trusting that the model's own representation handling matches what a purpose-built encoder provided. The chain-compression approach promises measurable speed and memory wins today while keeping the existing architecture intact. One is a bet on the direction the field is heading; the other is a bet on the hardware you already own.

Both bets are reasonable, and the labs pursuing them are not competing for the same deployment. PixelUMM's unified framing suits teams that want one model for understanding and generation. Sana suits teams that need high-resolution output on constrained hardware. The overlap is the part of the stack both are trying to delete.

NVIDIA has not said whether PixelUMM is finished. Sana's code is out, and the repo has grown to include sprint variants, video generation, ControlNet, LoRA, quantization, ComfyUI support and an online service. Two research teams, two routes, one target: the parts of the visual stack that nobody wanted to think about are being designed out.

Related articles