One Framework for Language and Vision: Horizon's 1.6B Open Model

A paper from Huazhong University of Science and Technology, Beijing Jiaotong University and Horizon Robotics proposes something that sounds almost too basic to be novel: a single model that treats language and images as the same kind of thing.
The framework is called Multimodal Flow, or MF-1, and it is fully open. Its largest version has 1.6 billion parameters, trained on 150 billion pre-training tokens. On GenEval it scores 82.8 and on DPG-Bench 75.3, putting a model of modest size alongside much larger multimodal systems.
The idea in one sentence
Most multimodal models are assemblies. A text model and an image model, or a language backbone with separate vision components bolted on, passing representations back and forth through discrete tokens or separate encoders. MF-1 does not do that. It models text and images together in a shared embedding space, and uses Flow Matching as a single generative objective and sampling procedure for both.

That is the interesting claim. Instead of translating between modalities at a boundary, the model operates on one continuous representation where language and vision are not fundamentally different objects. The authors argue this sidesteps the weaknesses of the discrete and hybrid approaches that dominate current practice.
If it holds up, the payoff is generality. A single framework that handles language modeling, visual understanding, image generation, editing, and video sequences is a much simpler thing to train, extend and reason about than a pipeline of specialists. It also points at a future where other modalities, any input that can be expressed as an ordered continuous block, fit into the same structure without a redesign.
There is a practical reason to want this beyond elegance. Every time you bolt a new modality onto a system, you introduce a new interface to maintain and a new place for errors to accumulate. Chaining a vision encoder to a language model means an image has to be summarized into tokens the language model understands, and that summarization loses information. A shared representation skips the translation step, which is where a lot of subtle quality leaks out of multimodal systems.
The name in the credits
One detail in the author list is worth a second look. Among the authors are Yu Kai, the founder of Horizon Robotics and formerly an architect of Baidu's deep learning efforts, and co-founder Huang Chang. The corresponding author is Wang Xinggang of Huazhong University's Vision Lab.
Yu's name on a paper is not routine. He has published rarely in the years since 2016, when he was part of an earlier attempt to unify text detection in natural scenes. His reappearance now, on a paper about unifying language and vision, reads as a deliberate return to a problem he worked on a decade ago, at a moment when the industry has the compute and the data to make it matter.
There is a telling line attributed to him from an earlier talk: that the buzzwords, VLA, VLM, end-to-end, world model, are often more meaningful commercially than technically, and that serious teams are largely doing similar things. A paper that collapses several of those categories into one generative framework is consistent with that view. It is an argument that the labels are less important than the underlying mechanism.
Why small and open is the point
A 1.6 billion parameter model that scores competitively is notable for the same reason the small models coming out of other labs are notable. Capability per parameter is the metric that decides what runs locally, what runs on a phone, and what a small team can afford to fine-tune.
The 150 billion training tokens are also modest by frontier standards, which suggests the approach is efficient rather than brute-forced. Whether that efficiency survives at larger scale is the open question. A small model that behaves well says something about a method, but not everything.
Open weights matter here for a specific reason. A unified generative framework is most useful if other people can extend it, test it on modalities the authors did not try, and build on the shared representation. A closed version would be an interesting paper. An open one is a tool.
What the benchmarks do not cover
GenEval and DPG-Bench measure how faithfully a model turns a prompt into an image. They say nothing about whether the unified representation improves a model's ability to reason about text and images together, which is the actual claim. A model could score well on image fidelity while treating the two modalities as separate subsystems that happen to share a number format. The authors are aware of this gap, and the paper's real argument lives in the architecture, not the scores. For anyone evaluating the work, the right question is whether the shared representation changes behavior on tasks that require both modalities at once, and that is exactly what the benchmarks leave unmeasured.
The honest unknowns
Unifying modalities in one space is a strong claim, and strong claims invite scrutiny. The benchmark scores are real, but they measure image generation fidelity, not whether the shared representation actually helps the model reason across modalities the way the framing implies. It is possible to get a good GenEval number and still have a model that treats text and images as neighbors in an embedding rather than as genuinely the same material.
The other unknown is scale. Continuous unified representations have historically been attractive in theory and stubborn in practice, because the training signal can be harder to stabilize than in a discrete setup. A 1.6B success is encouraging and not yet conclusive.
Where it points next
The obvious extension is video and audio, both of which are continuous signals and therefore natural fits for a framework built on continuous representations. The authors gesture at this, framing any modality expressible as an ordered continuous block as fair game. If the approach holds, the payoff is a single pipeline a team can extend to a modality the original authors never touched. That is the promise of open weights here: not a finished product, but a foundation others can bend to their own problem.
The bigger picture
What makes this worth watching is less the model than the direction. The field has spent years building increasingly elaborate bridges between separate text and vision systems. MF-1 is a bet that the bridges are unnecessary, that the right move is to stop treating language and images as foreign to each other and model them on one footing.
If that bet is right, the practical consequence is boring in the best way. One framework, one training objective, one set of tools, applied to whatever modality you need next. If it is wrong, it is still a well-run experiment from a team with a long history in the area, released in the open where it can be checked.
Either way, the interesting part is that the people who have watched this problem for a decade chose now to attack it again, and chose to do it with an open, small model rather than a giant one. That combination is usually a signal that someone believes the method, not the scale, is the thing that has been missing.
Related articles
A Wheeled Semi-Humanoid Finished an Hour of Laundry Without Help
Individual tasks can succeed while a workflow still fails. Dyna changed the metric.
LTX 2.5 Wants to Render Your Blocky Blender Draft Into a Finished Shot
You do not control what happens in text-to-video. This tries to fix that.
ServiceNow Turns Agent Failures Into Training Data
Generation without verification is noise. The gates are the product.
A Photonic Memory Wall Is the $88 Million Bet Volantis Just Made
The tradeoff everyone has learned to live with is the thing it is trying to delete.