← Back to blog
AiAbout 6 min read

Xiaomi Shipped an Omnimodal Model That Ties the Frontier, Plus the Training Recipe

Published Oct 4, 2026
Xiaomi Shipped an Omnimodal Model That Ties the Frontier, Plus the Training Recipe

A phone company released an AI model on 22 September and the headline number was 46. That is Xiaomi's score for MiMo-V2.6 Pro on the Artificial Analysis Intelligence Index, which the company says is the highest recorded by an open-source model at the time of release.

MiMo-V2.6 comes in two versions. Pro is the large one, Flash is the fast one. Xiaomi says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks, and points to improvements in coding, computer use, 3D reasoning and creative tasks. The models are omnimodal, meaning they accept text, images and other inputs through one set of weights, and they were trained with scaled reinforcement learning, which is the part that explains both the capabilities and the release format.

What actually came in the download

Most open-weight releases hand over the weights and stop there. MiMo-V2.6 ships weights, a technical report, the reinforcement learning environments, and the training code. Xiaomi published all four on its own site rather than routing developers through a model hub with a licence page and a waitlist.

The inclusion of RL environments is the meaningful difference. Weights let you run a model. Environments let you retrain it. When a lab releases the environments it trained on, it is telling you how it got the behaviour rather than only what the behaviour is. For anyone who wants to reproduce the result, or fine-tune the model against their own reward signal, the environments are the actual artefact.

It also raises the cost of the release for the company that made it. Publishing environments exposes the shape of the training signal, which is closer to a recipe than to a checkpoint. A competitor can read it and skip a year of trial and error. Xiaomi appears to have decided that the recruiting and credibility benefits outweigh that.

Why a device company is doing frontier research

Xiaomi is not a research lab that happens to sell phones. It is the other way around, and that orientation shows up in what an omnimodal model is good for.

A phone is the natural home for an assistant that can read a screen, understand a photograph, operate an interface and answer in a spoken conversation. Computer use, 3D reasoning and creative generation are not abstract benchmarks from that vantage point. They are the components of a system that has to run on a device someone is holding, often with a slow connection and a battery to protect. A model family split into a Pro tier for capability and a Flash tier for latency is a device-roadmap decision as much as a research one.

Releasing the weights serves two purposes. It recruits researchers who will improve the model in public, and it puts a floor under the company's position in a market where capability claims are otherwise taken on trust. A model you can download and evaluate yourself needs less marketing, and Xiaomi is competing for developer attention against labs whose entire reputation rests on published results.

What omnimodal means in practice

The label covers a specific engineering claim: one model handles several kinds of input through a shared representation rather than routing each input type to a separate specialist. For a phone, that matters because the alternative is running several models and stitching their outputs, and memory and latency are both scarce on a handset. Xiaomi's Flash tier exists for the same reason. A model that answers in a second on a flagship can be unusable on a mid-range device or in a car, so the family is really two points on a capability and latency curve rather than one release.

A clear glass lens with rainbow refraction resting on a dark ring

The training method explains the rest. Scaled reinforcement learning means the model is optimised against a reward signal rather than only trained to predict the next token, and it is the approach behind the recent jump in agentic performance across the industry. It is also the reason the environments matter as much as the weights. A reward signal is only as good as the environment that produces it, so publishing the environment is closer to publishing a method than publishing a result.

The device is the distribution channel

Xiaomi's advantage is that it does not need a developer ecosystem to reach users. It ships hardware to tens of millions of people, and a model that runs inside a phone assistant reaches more people in a quarter than a model hub reaches in a year. That is why the release format is generous. The weights and the recipe cost the company relatively little in lost licence revenue, and they buy credibility with the research community that would otherwise judge the model from a benchmark screenshot.

The risk is the same as the advantage. A phone assistant is judged on whether it works every time, and an agent that fails on one task in twenty is an annoyance in a chat window and a liability inside a device that also handles payments, photographs and messages. Capability benchmarks do not measure that gap, and a composite score of 46 does not either.

The same Tuesday, three other releases

The timing is instructive. On 22 September, Xiaomi's model appeared alongside Alibaba's Qwen-Image-2.1, which Alibaba open-weighted on 20 September and presented at its Yunqi conference the same day, and Tencent's Hy Image 3.5 preview, a professional image model priced at 0.15 yuan per 2K image on the Tencent Cloud API. Unitree used the same window to announce Dex5-S, a dexterous robotic hand with 22 degrees of freedom starting at 6,500 dollars.

Alibaba also used the conference to claim first place on Artificial Analysis' agentic leaderboard for Qwen3.8-Max and first on CodeArena for frontend programming, and to say that Qwen4 is in training with successor models planned at five to ten trillion parameters. Its head of Qwen-Image noted that the team's goal is to reduce the cost of completing a task rather than to maximize parameter count, which is the same trade MiMo's two-tier structure is making.

None of those announcements were coordinated and all of them pointed the same way. The competitive question in Chinese AI this autumn is not whether a frontier-class model exists. It is how much of one you are allowed to hold.

The benchmark caveat

A number like 46 depends entirely on the index behind it, and Artificial Analysis' Intelligence Index is one composite among several. The comparison to Claude Opus 5 and GPT-5.6 Sol is made by Xiaomi, on Xiaomi's choice of benchmarks, and the phrase "most agent benchmarks" is doing work that a table of results would do better. Agent benchmarks in particular are sensitive to scaffolding: the same model can score very differently depending on what surrounds it.

What makes the claim more than marketing is the download. Anyone who doubts the 46 can run the model, read the report, and retrain it in the environments Xiaomi shipped. That is a harder test than a leaderboard screenshot, and it is only available because the company chose to publish the recipe along with the result.

Related articles