Ant Group's Ling 3.1 Flash Activates 25B of 560B Parameters and Lands at 41

--- title: Ant Group's Ling 3.1 Flash Activates 25B of 560B Parameters and Lands at 41 meta_title: Ling 3.1 Flash Puts 560B Parameters Behind a 25B Bill meta_description: Ant Group's Ling 3.1 Flash scores 41 on the Intelligence Index with 25B active parameters out of 560B total, showing how far sparse MoE can stretch. ---
Ant Group's AntLingAGI released Ling 3.1 Flash on October 6. The number that stands out is the 25 billion parameters that actually fire when the model runs, out of 560 billion in total.
The model scored 41 on the Artificial Analysis Intelligence Index, a clear step up from its predecessor. The evaluation singled out agentic capabilities as the area of sharpest improvement, which is the category enterprise buyers have started to weight most heavily.
The architecture is the argument
Ling 3.1 Flash is a mixture-of-experts model. It carries 560 billion parameters in total but activates roughly 25 billion per token, and it supports a context window of up to one million tokens.
That ratio is the whole pitch. A dense model of comparable capability would need to move every parameter for every token generated. By routing each token through a small subset of experts, a sparse model delivers similar output while touching far less compute. The tradeoff is memory: all 560 billion parameters still have to be resident somewhere, even though most sit idle on any given forward pass.

The practical consequence is a model that can be served at a fraction of the cost of a dense equivalent at the same quality level, provided a team has the hardware to hold the weights. It is the same engineering story that has driven most of the frontier open-weight releases this year, and it is why the interesting metric has moved from total parameters to active parameters.
What the benchmarks cover
Ant Group published results across a range of evaluations: GDPVal-AA v2.1 at 1673 Elo, FrontierSWE at 75.16, HealthBench at 65.35, and Design Arena, where it ranks near the top for mobile applications and SVG generation.
The Intelligence Index score of 41 is the headline, and it puts the model in the company of releases from much larger labs. Ant Group says the full model will be open-sourced, with documentation already published. A two-week free trial is running in the meantime.
The engineering demonstrations are more telling than the scores
Ant Group also described three worked examples, and these say more about intended use than a leaderboard does.
The first: building a Lua-to-x86 compiler from scratch, passing 178 of 182 tests. That is a task where the model has to hold a large specification in mind, produce internally consistent code across many files, and keep track of edge cases throughout. The work punishes models with weak long-context handling.
The second: accelerating type checking on a large Python codebase. This is a maintenance task rather than a greenfield one, which is where most real engineering time actually goes. A model that can reason about an existing repository's type relationships is useful in a way a model that writes fresh functions is not.
The third: building a desktop agent that coordinates a browser, local files, and terminal tools. Multi-tool orchestration is the current frontier of agentic work, and it is the capability the benchmark jump points at.
All three are company-reported and have not been independently reproduced. The 182-test compiler claim in particular would need outside verification, since passing tests is only meaningful if the tests are representative. A compiler that handles common syntax and fails on the unusual cases can still score well on a suite that was written the same week as the model.
There is a broader pattern to note here as well. Chinese labs have increasingly used engineering demonstrations rather than benchmark scores as their primary marketing, because benchmarks have become ambiguous and demonstrations are concrete. The tradeoff is that a demonstration shows one trajectory through the problem space, and nothing about the distribution of outcomes.
The longer arc: what Chinese open weights keep doing
Ling 3.1 Flash fits a pattern that has held all year. Chinese labs release frontier-scale open weights with permissive licenses at a pace no Western lab has matched, and they compete on efficiency rather than raw scale.
The context is worth stating plainly. A developer running one of these models owns the inference stack, which means no per-token API cost, no data leaving the building, and no third party deciding when the model changes. For teams with privacy constraints or predictable-but-heavy workloads, that is worth real money.
Victor Taelin noted this month that no open model remains in the top 25 of the Artificial Analysis rankings, with the best, Xiaomi's MiMo, sitting at 26. Ling 3.1 Flash at 41 on the Intelligence Index is an intelligence score rather than a rank, so the two figures measure different things, and the gap between the best open model and the best closed model is still real. What has changed is that the gap is now measured in points rather than in whether an open model can do the job at all.
Where this lands for anyone choosing a model
Three considerations matter more than the parameter ratio.
License and deployment terms. Ant Group says the model will be open, but the specific license governs whether commercial use is unrestricted. Given how often "open weights" and "permissive license" get used interchangeably when they are not the same thing, the terms deserve a careful read before any production dependency. The same caution applies to the documentation published ahead of the release, which describes capabilities rather than obligations.
The context window claim. One million tokens is a large number, but usable context is usually shorter than advertised. Whether the model retains quality across the full window, or degrades well before it, is the question that determines usefulness for long-document and large-repository work. Ant Group's own long-context demonstrations suggest the model handles large inputs, though again these are vendor tests.
The memory requirement. A 560B sparse model is not a laptop deployment. The active-parameter story reduces compute per token, not the footprint of the weights, and anyone planning to self-host should size the hardware to the total, not the active count. That distinction trips people up regularly. A model described as activating 25B parameters sounds small until you try to load it, at which point the full 560B has to fit somewhere.
There is also a practical support question. A model from a lab that ships frequently will be superseded, and teams should plan for a migration path. Ant Group's release cadence suggests Ling 3.2 or Ling 4 is a matter of months, which is good for capability and awkward for anyone who built a fine-tune or a deployment script around one version.
Ling 3.1 Flash is a capable release from a lab that has been shipping steadily. The real news is structural: a 25B-active model can now credibly sit in the same conversation as much larger dense systems. That is the argument the entire sparse-MoE line of work has been making, and the benchmarks are finally starting to support it.
Related articles
The Gap Between Arena Leaderboards and Real Image Output Is Getting Wider
The infrastructure for ranking models has never been better, and the connection between rank and practical output has never been looser.
Google Flow and Adobe Firefly Move AI Video Out of the Chat Box
Base model quality has converged enough that the differentiator has moved to what surrounds the model.
Google Cut Nano Banana 2.1's Output Price in Half and Fixed Its Weakest Features
The most consequential detail sits outside the feature list, and it is the price.
Vida Wants To Bill for AI Agents by Results Rather Than Usage
Usage-based billing aligns the vendor's revenue with the agent taking longer. Outcome pricing inverts that.