DeepSeek Open-Sources the Engineering Experience Behind Running Trillion-Parameter Models: This Time, Domestic Compute Is Filling In the Software Layer

On September 30, DeepSeek released a set of code on its official WeChat account, open-sourcing six modules for Huawei's Ascend platform: TileLang, DeepGEMM-Ascend, DeepEP-Ascend, TileKernels, FlashMLA, and DeepSelect. The names sound scattered, but what they cover is actually quite coherent: a programming language, a batch of compute kernels, and distributed communication.
This doesn't look like a routine open-source release. What it open-sources isn't a model, but the foundation layer needed to get a model running.
Once the Chips Are Good Enough, Something Else Becomes the Bottleneck
In recent years, domestic compute has mostly been discussed in terms of raw compute figures. The spec sheets look better every year, but any engineer who has actually migrated a model knows the trouble rarely comes from peak compute.
NVIDIA has spent more than twenty years building the CUDA ecosystem. Compilers, math libraries, communication libraries — layer piled on layer, so developers can use a card the moment they get it, with documentation, community, and troubleshooting notes all readily available. Switch to another hardware stack, and the same operator may have to be written from scratch, the scheduling tuned by hand, and the multi-node, multi-GPU communication libraries implemented yourself against the spec. The chip's benchmark scores may not be low, but the act of moving a model onto it is absurdly expensive.
There's another layer to this gap that rarely gets mentioned. Hardware can be bought with a one-time investment; a software ecosystem can only be ground out by people, year after year. Buying cards is a procurement problem; completing the toolchain is a matter of time — and that time can only be traded through real projects. Whichever team gets its model running, running stably, and running at scale on a given platform first is the one that can accumulate reusable experience. That's also why this open-source release deserves to be discussed on its own: it isn't about what was bought, but about someone handing over what they accumulated.
So domestic compute's real, long-term weakness is software. The chips can compute, but they aren't easy to use. What DeepSeek open-sourced this time lands exactly in that gap.
Among the Six Modules, TileLang Is the One Worth Watching Most
Among the six modules, the heaviest is TileLang. This is DeepSeek's self-developed high-level programming language, previously aimed mainly at the NVIDIA backend; this time it officially supports Ascend 950, offering native code generation, automatic scheduling, and synchronization. DeepSeek's positioning of it is blunt: an alternative to CUDA, and with a "simpler programming model."
The accompanying TileKernels solves another long-standing problem. It can automatically select the NVIDIA or Huawei backend, exposing the same set of Python APIs to the layer above. In other words, the same code runs on two hardware platforms, and the cost of switching is reduced to the configuration level rather than rewriting the kernel.
The internal benchmark data made public isn't vague either. Some kernels reached a very high proportion of the nominal hardware ceiling on Ascend 950DT: BF16 dense matrix multiplication at 99.8%, FP8 at about 99.5%; the sparse attention implementation targeting DeepSeek V4.1 ran at 410 TFLOPS in the prefill stage, roughly 95% of the theoretical value. The two companies also jointly optimized a supernode design based on 128 Ascend 950 chips, specifically to balance computation and data movement.
The value of these numbers is in proving that this path works. In the past, when people talked about domestic compute, they talked about whether it could be used at all; now teams are starting to be willing to publish their experience squeezing performance out of the hardware, which shows that at least some have pushed it far enough to share.
One Set of APIs Hanging Off Two Backends: Where's the Value?
Looking at TileLang and TileKernels together makes their use clearer. TileKernels can automatically choose the NVIDIA or Huawei backend, exposing the same Python APIs upward. That means a team maintains one codebase, not two.
In practice, a reasonably serious inference service usually has to run on several kinds of hardware at once. The cloud may still be NVIDIA, self-hosted or domestic-substitution scenarios use domestic cards, and edge devices are yet another stack. If every hardware change requires rewriting the kernels, a team has to maintain several parallel implementations and test each one separately. A single set of APIs downgrades this from "rewrite" to "change the config," saving not just labor but also chances to make mistakes.

For a team that just wants to ship a product, this kind of change won't make it into any release note, but it determines whether they're willing to try domestic hardware.
The Real Change Happens in Migration Cost
For an ordinary developer, the most practical impact of this open-source release is migration cost.
In the past, migrating training or inference to Ascend meant filling in low-level operators yourself and adapting the tools yourself; it wasn't rare for a project to spend months on it. Now these toolchains are open-sourced directly, and many pitfalls have already been hit and fixed by others and can be reused as is. The bar drops from "you need a dedicated team for operator adaptation" to "someone on the team needs to understand this toolset."
That also explains why this release drew a sizable reaction in developer communities. Plenty of models have come out lately; what's truly scarce is skipping a stretch of the road.
What to Watch Next Is Whether Others Use It
To judge the weight of this open-source release, you can't just look at what DeepSeek itself says; you have to look at what others do over the next six months.
Whether a toolchain can stand on its own depends on whether third parties are willing to invest in it. If TileLang only serves DeepSeek's own models, its significance stays inside that company; if other teams use it to write operators, file issues, and extend hardware backends outward, it starts to become public infrastructure. No matter how pretty the numbers are, they only count once they've been tested by months of real workloads.
Another variable is whether other domestic chips will follow and plug in. Currently the public information centers on Ascend; if this high-level language can be migrated to more platforms along the same logic, its value would rise another notch. There's no answer to that yet.
Open-Sourcing the Foundation Is Slower Than Open-Sourcing Models
Over the past two years, the main battlefield of domestic open source has been model weights. Parameters, leaderboards, download counts — all things you can see immediately. The underlying toolchain is different: it doesn't appear on leaderboards, and it doesn't generate buzz in the short term. But without this layer, no matter how many models are stacked on top, it all hangs in the air.
By handing over its engineering experience running trillion-parameter-scale models on Ascend, DeepSeek is doing exactly this kind of unglamorous make-up work that determines whether the industry can move forward.
Whether to migrate the NVIDIA projects you have on hand to Ascend right now still depends on the specific scenario, and on whether anyone on the team has time to spend at the low level. But at least from this day on, the excuse that "the ecosystem isn't usable" is starting to hold less water. And once an ecosystem starts to be usable, what happens next often comes faster than expected.
Related articles
13,000 Internal Screenshots Ended Up on Public GitHub, and No Attacker Put Them There
A default behaviour, repeated across a fleet, is a policy outcome.
Amazon Wants Investors to Own $8 Billion of Nvidia Chips It Still Uses
Airlines have leased back planes for decades. Now the same idea is being applied to GPUs.
OpenAI Traced a Reasoning-Extraction Campaign to People Tied to Moonshot AI
The model became the decryption oracle for its own hidden reasoning.
The First AI Film Festival Paid Out $450,000 and Taught a Lesson About Story
The winning films used the tools to serve an idea that already existed.