China's New Open Models Are Being Trained to Be Agent Substrates

Three model releases out of Chinese labs in the last week share a design instinct that is easy to miss if you read them as a list of parameter counts. None of them was built to be good at conversation.
They were built to be the thing an agent runs on. That is a different target, and the training choices that follow from it look strange if you assume the goal is a better chatbot.
Training the task and the world around it
The clearest example is IQuest-Q1, released by the ZhiZhi Innovation Research Institute on September 29. It is a sparse mixture-of-experts model with 320 billion total parameters, 15 billion active, and a 524,288-token context window. The architecture is unremarkable next to its training method.
Rather than fine-tuning on transcripts of human conversation or on static instruction sets, the team synthesized tasks and the environments they happen in together. The model then trained across several different agent scaffolds. It kept each scaffold's native tools and context management intact, and treated only the model's own mistakes as learning signal. The harness was not allowed to become the lesson. Afterward, four specialist models were merged into one through multi-teacher on-policy distillation.
This matters because of a problem anyone building agents has run into. A model that scores well on a benchmark can fall apart when you wrap it in your own tools, because it learned the quirks of somebody else's harness. Training across scaffolds while holding the harness constant attacks that specific failure. The reported result is a model that lands just behind Claude Opus 5 on natural-language-to-repository tasks, which is a strange thing to be good at and exactly what a coding agent needs.
Seeing and deciding in the same step
The second release takes a different route to the same goal. On the same day, the Shanghai AI Laboratory introduced a decision model called Shusheng Mingjue, in three sizes: 0.8B, 2B, and 4B. It is small on purpose.
The design principle is that perception and decision should not be separate stages. Most agent stacks take a screenshot, hand it to a vision model, get a description, pass the description to a planner, and act. Each hop adds latency and loses detail. Mingjue fuses native visual perception with instruction following so the model looks at a screen and makes a call in one pass. The lab reports it beating Jev and comparable open models on decision quality. For an agent that has to act inside a dynamic environment, the size is the point: a 4B model can run close to the loop, and a 320B one cannot.
Deleting the attention layer to reach a million tokens
The third release is the most architecturally aggressive. Naive N0.5 Flash, from the Beijing lab NaiveAI, is a 309-billion-parameter mixture-of-experts model with 15.5 billion active parameters, built on Xiaomi's MiMo-V2.5. It natively supports a one-million-token context and ships under the MIT license.
The interesting part is what it removed. The model uses a hybrid of sliding-window attention and DeepSeek's sparse attention, and it has no full-attention layers at all. Standard transformer attention scales quadratically with sequence length, which is why long context has historically been expensive. Removing full attention is a bet that the useful information in a long sequence can be reached through structured sparsity, and if that holds up it changes what an agent can keep in view at once. A million tokens is roughly a large repository, or a long working session's worth of history, held without truncation.
The smaller releases point the same way
The pattern shows up beyond these three. Tsinghua's Puro-2B was trained for around $4,400 and beats a larger Qwen model across fifteen tasks on average. A telecom lab shipped TeleOCR, a 1.2B document-parsing model that topped a document-understanding benchmark. Stanford and NVIDIA released a contrastive verifier that picks the best candidate action by embedding alignment rather than by reasoning, and reports large speedups over a comparable judge model.
Every one of those is a component in an agent, not a destination for a user. A verifier that cheaply ranks candidate actions, a small model that parses documents, a tiny model that decides what to click. The pieces are getting specialized and getting small, while the one model that holds the whole session's context gets very large.
The evaluation problem nobody has solved
There is a hole running through all of this. Benchmarks for agent base models are still mostly borrowed from chatbot evaluation, which measures the wrong things. A model can score well on a reasoning test and still be a poor substrate for an agent, because the properties that matter only show up across a session: how many turns pass before the context degrades, whether a tool call that fails gets retried sensibly, whether the model notices that it has lost track of the original goal.
Most of the reported numbers here come from the labs themselves, on benchmarks they helped define. IQuest-Q1 "just behind Claude Opus 5" on natural-language-to-repository tasks is a vendor comparison. Mingjue "beating Jev" is a vendor claim. Naive N0.5 Flash's absence of full attention is an architectural fact that is easy to verify and whose practical consequences are not yet measured at scale. None of that makes the direction wrong. It means the honest state of play is that we have three interesting designs and very little independent evidence about which one holds up when an agent runs for six hours.
The gap is starting to be acknowledged. Independent groups have been building benchmarks aimed at production serving rather than code generation, and at agent hijacking rather than model safety, which suggests the field knows it has been measuring the wrong layer. Until those mature, the useful signal from a release like this is not the score. It is the training method, because a method that addresses a known failure mode of agents is more likely to be real than a number that addresses a leaderboard.
There is one more consequence of the shift that is easy to overlook. If the important models become small components rather than large destinations, then a lab's advantage stops coming from having the biggest model and starts coming from knowing how the pieces fit together. A team that can train a fast decision model, a cheap document parser, and a verifier, and wire them into a loop that runs on modest hardware, has something a single giant checkpoint cannot match. That is a different kind of competition, closer to systems engineering than to scaling, and the releases landing this month suggest at least a few Chinese teams have already reoriented toward it.
For two years, the open-weight race was measured by how close a free model could get to a frontier chatbot. That framing is losing its grip. Chat quality has become a baseline, and the questions that decide whether software works have moved to different ground: how many turns before the context degrades, how fast the model can act inside a loop, how well it tolerates being dropped into someone else's tooling.
The models that answer those questions will not win many leaderboards. They will win the less visible contest of which one a developer reaches for when the chatbot has already done its job and the work needs to actually happen. That contest played out quietly this week, in three releases that were not trying to impress anyone with conversation.
Related articles
Agility Digit 5 Ships With a Safety Case, Not Just a Spec Sheet
A warehouse floor is not a lab. Certification is the gate, not the demo.
Frontier Agents Finished 30 Percent of a Research Workflow. That Is the Number.
Agents can run research. Inventing the procedure is still out of reach.
Figure AI Locked In $3.5 Billion of Compute Before It Has a Product to Sell
The bet is that generalisation is a compute problem. The field has not settled that.
OpenAI Finally Put Transparent Backgrounds in the Image API
A small feature that deletes a whole step from the pipeline.