Kimi K3 Enters OpenAI's Enterprise Billing System: The Overseas Path for Chinese Open-Source LLMs Is Being Rewritten

On October 1, the U.S. AI infrastructure company Baseten confirmed one thing: enterprise developers can directly call Moonshot AI's Kimi K3 inside OpenAI's programming tool Codex, and the resulting costs are charged directly against the enterprise's existing OpenAI procurement credits. There is no need to add a new vendor entity, nor to go through procurement approval again. This is the first time a Chinese open-source large model has entered OpenAI's enterprise billing and settlement system.
The weight of this matter lies not in parameters or leaderboard rankings, but in the fact that it changes enterprises' path dependency in procurement.
The Procurement Process Itself Is the Barrier
When a large enterprise onboards a new model vendor, the trouble is never technical integration, but compliance and financial processes. A new vendor must go through onboarding review, set up separate billing and settlement systems, and go through legal and procurement again. For a team that just wants to verify whether a model is any good, the cost of this process is often far more expensive than the model itself.
By settling through OpenAI credits, Kimi K3 bypasses this entire set of frictions. AI budgets already approved within an enterprise can be spent directly on Kimi K3, while finance still sees an OpenAI bill. As the underlying gateway, Baseten plugged Moonshot AI's model capabilities into OpenAI's ecosystem channel.
Before this, Kimi K3 had already landed on Amazon Web Services' Amazon Bedrock platform, and Moonshot AI thus became the first Chinese AI company to partner with a leading global cloud provider under a revenue-sharing model. One is revenue sharing on a cloud platform; the other is credit drawdown in a developer tool. Both paths point to the same result: Chinese large models are moving from the tech demo stage into scaled commercial deployment.
How Good Is the Model Itself?
Kimi K3 was released in July this year with 2.8 trillion parameters, making it the largest open-source large model in the world at the time of release. It supports a 1-million-token context window, natively has visual understanding capabilities, and has been specifically optimized for code development, deep research, and complex knowledge processing.
On the day of release, Kimi K3 topped Arena, an international code evaluation leaderboard, becoming the first Chinese large model to reach its top spot. Tesla CEO Elon Musk also publicly commented that the result was impressive.
More noteworthy is its architectural approach. Kimi K3 uses Kimi Delta Attention and Attention Residuals to improve information flow in long sequences, and adopts a Stable LatentMoE structure that activates only 16 of 896 experts. The company says this improves scaling efficiency by about 2.5x compared with the previous generation. This design direction is consistent with the collective moves of Chinese labs in recent months: not competing over whose parameters are larger, but over how much usable capability each unit of compute can deliver.
A 48-Hour Autonomous Chip Design
If entering Codex is a commercial matter, another development is closer to technology itself.
According to public information, Kimi K3 has completed a proof of concept for AI autonomously designing a chip. In a continuous 48-hour nonstop agent test, it relied on open-source EDA tools and a 45nm process library to independently complete the design, iteration, and full-process verification of an adapted chip. The output was a lightweight 4-square-millimeter chip with 1.46 million standard cells and 0.277MB SRAM, achieving stable timing convergence at 100MHz and maximum decode throughput of over 8,700 tokens per second.

Such results should be viewed cautiously. A proof of concept does not equal a mass-producible chip, and 45nm is not an advanced process node. But it validates a closed-loop workflow: whether a model can carry an engineering task that requires multiple rounds of trial and error from start to finish without step-by-step human intervention. For people building agents, that is more telling than whether a single image looks good.
Where Is the Competitive Layer Heading?
Kimi K3's overseas push is running in parallel with another line.
In the same week, Zhipu released the open-source model GLM-4.6, focusing on upgrading its agentic coding capabilities and completing adaptation to Chinese chips such as Cambricon and Moore Threads, enabling stable deployment of FP8+Int4 mixed-quantization inference on Chinese hardware. DeepSeek also open-sourced its desktop agent runtime framework Harness v0.2 on October 2, adopting an "everything is a plugin" architecture, with built-in capabilities for file organization, data analysis, document drafting, and code modification, released under the MIT license.
Putting these together, a direction becomes clear: the model itself is becoming infrastructure, and real differentiation is starting to shift to two places. One is the runtime framework—the layer beyond the model: whoever can let developers and enterprises integrate models into their workflows at low cost. The other is hardware adaptation: whoever can run inference on cheaper, more self-controllable compute.
Kimi K3 has already teased K3.1, which supports a 1-million-token context and offers three reasoning effort levels—Low, High, and Max—for users to choose based on task complexity. Long context and selectable reasoning levels are becoming standard features for Chinese open-source models competing for developers.
In the past, Chinese large models going global mostly involved one-to-one engagement with overseas customers; for large enterprises to adopt them, they had to add supplier qualification reviews, creating a high barrier to deployment. Kimi K3's two steps have lowered that barrier somewhat. As for whether enterprises are willing to pay for it, the next few quarters will provide the answer.
Related articles
AI Video Tools Rate 9 Out of 10 for Ease of Use. The Compliance Score Is a Different Story.
Users love AI video generators. Legal departments have not been asked yet.
InternLumina-U2 Puts Text, Images, Video, and 3D in One 16B Model, Then Ships Two Sets of Weights
The task list is ambitious. The release format carries more signal.
HappyOyster Sells a World You Can Walk Into, Not a Clip You Watch
Most video models hand you a file. This one hands you a room with a door in it.
Luma Ray3 Modify Keeps the Performance and Renegotiates the World Around It
The hardest problem in AI video editing is not the effect. It is not wrecking the take you already paid for.