← Back to blog
AiAbout 6 min read

Domestic Large Models Rush to Open Source in September: This Time, the Competition Is Over “Affordability”

Published Oct 2, 2026
Domestic Large Models Rush to Open Source in September: This Time, the Competition Is Over “Affordability”

September is only half over, and people doing evaluations are already struggling to keep up with the update cadence of domestic large models. New releases from Zhipu, Meituan, ModelBest, Tencent, and Ant Group have landed in almost the same time window, spanning general-purpose large models, on-device small models, image generation, and embodied intelligence. Beyond the buzz, what is truly worth examining is the path they chose in unison: instead of piling on parameters, they are pushing hard on “how to make models affordable to use and deploy.”

Let’s start with Zhipu. GLM-5.3-Flash, released in September, has 320 billion total parameters, but only 18 billion active parameters actually involved in computation each time. The sparse activation approach is straightforward: make the model larger, but wake up only a small fraction for each inference, bringing down both cost and latency. It is open-sourced under the MIT license, natively supports multimodality, and costs about $0.25 per million output tokens. The heavier statement in the release notes is this: the inference service runs entirely on domestic AI chips, end-to-end performance is about three times higher than the initial baseline, and it has reached a scale of 100 trillion free tokens per day. For developers, that sentence carries more weight than any number in the parameter table.

Over at Meituan, LongCat-2.0 has 1.6 trillion total parameters and natively supports a 1 million-token context. Its highlight is not size but the soil it was trained in: according to the company, it is the industry’s first trillion-parameter open-source large model to complete full-process training and inference on a 10,000-GPU-class domestic compute cluster. Alongside it are the AI business workbench CatPaw and the local-life agent “Xiaotuan,” the latter of which has cumulatively responded to over 700 million user requests. A company that started in local life putting a trillion-parameter model on stage shows that the cost structure of this track is loosening.

ModelBest’s MiniCPM5-2B goes the opposite direction, squeezing parameters down to the 2-billion level and aiming for on-device deployment. It continues MiniCPM’s long-standing approach: small parameters, practical deployment, able to work when installed on a phone or edge device. At almost the same time, MiniMax open-sourced Code CLI under the MIT license, and Zhipu added GLM-5.3-FlashX, focused on high throughput, with inference speed of about 200 tokens per second. With the on-device, inference-acceleration, and coding-tool lines all filled in together, this combination says more about how mature the ecosystem has become than any single “flagship.”

Putting these together, a trend becomes clear: open source is turning into a competitive weapon, not a matter of sentiment. MIT licenses, open weights, free quotas—these moves point to something very practical: whoever pulls developers and small and medium-sized enterprises into their ecosystem first will be first to capture the next round of API calls. A parameter lead can carry a launch event, but it cannot carry day-to-day production environments.

The image line was not absent either. On September 4, Bailing Lab under Ant Group released and open-sourced the LLaDA-Image series, including Base and Turbo versions, with about 7 billion parameters. A single checkpoint simultaneously supports text-to-image, reference-image editing, and Chinese/English text rendering, without the need for an additional editing branch. The highlight is the Turbo version: it uses Twin-DMD distillation to compress sampling steps to 2–4 steps, preserving both quality and speed, and inference does not depend on high-end GPUs. On Qwen-Image-Bench, it scored 53.53 in English scenarios and 53.38 in Chinese scenarios, which the company says are the current state of the art. Reducing steps from the usual 50 to single digits matters not only for speed, but also for making real-time generation viable on consumer-grade hardware.

On the embodied intelligence side, Zidong Taichu open-sourced ZDTaichu5.0-9B on September 15. This 9-billion-parameter model took eight global firsts in its parameter class for spatial understanding and embodied tasks, and scored 78.27 on MindCube-tiny, 20.67 percentage points higher than Qwen3.5-9B. What it fills in is the robot’s “what to do after seeing” segment: recognizing a cup only answers “what it is”; judging the orientation of the cup handle and whether something can be placed beside it gets closer to “what can be done.”

Then there is Tencent, which cannot be ignored. On September 22, Hunyuan released Image 3.5 Preview, focusing on improving text generation, composition, realism, and continuous editing, with support for uploading up to five reference images at a time and output up to 2K. Pricing is set at 0.15 yuan per 2K image, and only final outputs are charged. The release date was also carefully chosen: it landed exactly on the opening day of Alibaba Cloud’s Apsara Conference. This kind of “timed release” is no longer new in China, but for image models, it effectively drags competition from the technical level straight to attention and ecosystem entry points.

Stringing these moves together, September’s main thread can be summarized in three things: faster open-sourcing, cheaper inference, and autonomous compute. The three actually interlock. Open source lowers the barrier to use; low-cost inference turns that barrier into real usage volume; and domestic compute substitution determines whether this low pricing can be sustained. GLM-5.3-Flash dares to use MIT plus free quotas; LongCat-2.0 dares to run a trillion parameters on a domestic 10,000-card cluster. Behind both is the same judgment: only when costs are controllable do you dare to hand the model over.

For ordinary creators and SMEs, the direct impact of this round of changes is that the math finally works. In the past, if you wanted to use frontier models, you either had to bear a not-inexpensive API bill or piece together your own GPUs—two barriers that kept most people out. Now, 0.15 yuan for a 2K image, real-time image generation in 2–4 steps, and on-device models that run with just 2 billion parameters have pushed these costs into a range where many people will actually give it a try. When tools get cheaper, trial and error increases, and quality is often honed through exactly this repeated trial and error.

An open vault door revealing server racks and a glowing microchip, representing open-source AI models on domestic silicon

There are also places where we should stay clear-headed. The fact that open-source licenses are becoming more permissive does not mean commercial use comes without cost: non-commercial clauses, data compliance, and controllability of model behavior all still need to be verified one by one in real business. The threefold inference performance improvement on domestic chips is based on each party’s own benchmarks, and third-party replication has not caught up; cases still exist where a model is “open-sourced” but its training code is not public, which limits reproduction and secondary development. Good-looking numbers are one thing; stable delivery is another.

In September’s dense wave of updates, the real signal is not “domestic models have set a few more firsts,” but that the cost logic across the entire supply side is being rearranged. As models themselves look more and more like infrastructure, competition will move downstream: whose workflow is smoother, who captures scenarios more accurately, who can turn cheap compute into consistent output. Open source has only fired the starting gun; there is still a long way to run.

Related articles