Alibaba Puts Its LLMs on Its Own Silicon: Zhenwu V900 Triples Performance, Qwen 4 Already in Training

At the Apsara Conference in Hangzhou in late September, Alibaba put three things on the same table: a new chip, a model in training, and a data center bill running through 2032. In past years, models usually took center stage at this kind of event; this year, what got mentioned over and over was silicon.
Alibaba's T-Head unveiled the Zhenwu V900, an AI chip that handles both training and inference. According to the official line, it delivers three times the compute of the Zhenwu M890, with 216GB of memory, 1200GB/s of inter-chip interconnect bandwidth, and native support for both FP8 and FP4 low-precision formats. The Zhenwu M890 was only released in May this year, just four months earlier. The V900 won't go into mass production and go on sale until the first quarter of 2027, so what's on stage now is still a spec sheet, not inventory you can order.
What's really worth looking at is the networking behind it. T-Head paired the V900 with its in-house ICNSwitch interconnect chip, built around memory semantics and unified memory addressing, and claims that over a thousand V900s can work together like a single "superchip." The supernode servers launched alongside it pack in, besides the V900, the company's own smart NICs and SSD controller chips — compute, storage, and networking all held in its own hands. By Alibaba's disclosed figures, a single cluster can scale to 500,000 cards.

Alibaba already has a batch of real workloads in hand. Supernodes based on the Zhenwu M890 have previously run models with more than two trillion parameters, such as Qwen3.8 and Kimi K3, and Zhenwu-series chips have served more than 650 enterprise customers in total, spanning industries including autonomous driving, finance, energy, and manufacturing. Those numbers explain why Alibaba dares to call for the next-generation cluster to reach 500,000 cards.
The moves on the model side are more direct. Alibaba confirmed that Qwen 4 is already in training, and that Qwen 4.5 and Qwen 5 target parameter sizes between 5 trillion and 10 trillion. For reference, Alibaba's current strongest model, Qwen3.8-Max, has 2.4 trillion parameters. Bigger parameters and bigger clusters are two sides of the same thing: without enough cards, training cycles for trillion-scale models stretch out until they lose their meaning.
Alibaba Cloud has set an infrastructure target of more than 20GW of global data center capacity by 2032, and plans to open three new regions — Turkey, Finland, and the Netherlands — in the coming year, while expanding capacity in Malaysia, Germany, the UAE, France, and elsewhere. Telling the chip, model, and cloud story as one strategy was the clearest throughline of the conference.
Impressive numbers, but the baseline is its own chip
When judging this roadmap, there are two things to watch in how the numbers are framed.
First, the threefold improvement compares the V900 with its own M890 — not with Nvidia's Blackwell, and not with any third-party accelerator card. Memory capacity and interconnect bandwidth don't translate directly into real training throughput, inference latency, and cost per token. Alibaba also didn't publish benchmark scores from independent third-party testing.
Second, the software ecosystem. Mainstream large-model training code worldwide has long been written around CUDA, and the driver libraries, containers, and orchestration tools all rest on years of accumulated work. Moving a workload that runs on Nvidia's platform to a new chip often costs you not in hardware specs but in the weeks engineers spend rewriting code and tuning performance. Alibaba didn't explain how hard migration between the V900 and the existing Nvidia toolchain would be, nor did it publish pricing, available instance types, or regional support.
This isn't Alibaba's problem alone. Capacity at advanced process nodes, yields, and performance per watt are all real constraints that domestically developed chips have to face. The significance of the Zhenwu V900 lies more in pushing the question of "can we build it ourselves" one step forward than in matching anyone on any single metric.
Why now
The timing deserves a closer look. Over the past two years, U.S. restrictions on exports of advanced GPUs to China have kept tightening, forcing domestic cloud providers to invest faster in their own accelerators. Alibaba and Huawei, plus smaller players like Enflame and Iluvatar CoreX, are all pushing in the same direction. During the same period, reports said Huawei was also moving up the launch timeline for its next-generation AI chip. Both companies say roughly the same thing publicly: replace Nvidia's position in the Chinese market before the rules clamp down further.
For Alibaba itself, in-house silicon pays off in two ways. One is reducing its exposure to swings in U.S. policy, so compute planning no longer rises and falls with export rules; the other is keeping in-house the chip-plus-cloud profits that Nvidia historically captured.
If you only look at model releases, this round of competition still looks like a contest over who has more parameters and higher benchmark scores. Look at the chip, the supernode, the cloud regions, and the model roadmap together, and the battlefield has actually shifted down to the layer of silicon and machine rooms. Models are the facade; whether you can build your own cards, wire up your own networks, and supply your own power is what determines how high the building can go.
The test comes later
Over the next twelve to eighteen months, three things will give the answers. Whether the 500,000-card supernode claim holds up in real customer deployments; whether Qwen 4.5 and Qwen 5 can really scale to close to 10 trillion parameters; and whether the 20GW data center target can secure the corresponding financing and electricity.
If even half of any one of these comes to pass, it would be enough to show that China's AI hardware stack is closing the gap with Nvidia. That narrowing may not show up in parity on any single metric, but more likely in deployment scale itself. Alibaba has said its piece in Hangzhou; the rest is up to the production lines and the electricity meters to answer.
Related articles
Salesforce Pays $2 Billion for a Company That Interviews Your Customers For You
Interviews are evidence. Digital twins are a prediction. The line between them is the test.
AMD Buys Fei-Fei Li's World Labs for $8.2 Billion to Own Physical AI
AMD is buying a research lab, and paying a chipmaker's price for it.
Google Put a TPU in Orbit and Started Counting the Cost of Space Data Centres
One working chip proves the trip is survivable. It says little about the profit.
Someone Catalogued 13,000 Ways AI Writing Gives Itself Away
The tells did not disappear. They moved somewhere harder to see.