← Back to blog
NewsAbout 6 min read

OpenAI Traced a Reasoning-Extraction Campaign to People Tied to Moonshot AI

Published Oct 4, 2026
OpenAI Traced a Reasoning-Extraction Campaign to People Tied to Moonshot AI

On September 30, OpenAI published an account of what it called a coordinated campaign to pull protected reasoning out of its models. The company said it disrupted the operation and attributed a core cluster of the activity to individuals associated with Moonshot AI, the developer of the Kimi model family.

The disclosure is short, and it is careful about the line between a company and a group of people. OpenAI did not accuse Moonshot AI of running the campaign. It said the people it identified were associated with the firm. That distinction is doing legal work, and it should be read as written.

What the campaign did

Distilling a model normally means training a smaller model on the outputs of a larger one, so the small model learns to imitate the big one at a fraction of the cost. That is common and often legitimate, which is why labs put effort into limiting it rather than banning it outright.

OpenAI described a more adversarial version. Operators manipulated interactions to get the model's protected reasoning reproduced in visible form. One technique involved copying encrypted reasoning from one conversation and asking a model in a different one to decrypt it. The reasoning traces that frontier models keep hidden are hidden for a reason: they expose how the model arrives at an answer, and they are the most profitable thing to steal.

The company described a fast escalation. The activity began on July 1 and spiked on July 24 and 25, with around 16,000 attempted extraction requests from more than 4,000 users. A related cluster involving more than 15,000 users was fully disrupted by July 28.

OpenAI said no encryption was broken and no databases were breached. It closed the reasoning-replay pathway, banned accounts, and shared its findings through the Frontier Model Forum and government channels.

Why the encrypted-reasoning angle matters

The interesting technical detail is the cross-conversation replay. If a lab encrypts the reasoning trace and keeps the key away from the user, the trace looks safe. The attack does not try to break the encryption. It moves the ciphertext into a fresh context where the model itself is asked, in effect, to make sense of it. The model becomes the decryption oracle.

That is a design lesson more than an incident. Any system that hands a model data it should not reveal, and then lets the same model answer questions about that data, has a soft spot. The fix is architectural: keep the protected content out of any context the model can be prompted to translate.

The economics underneath

Reasoning traces are valuable because training a competitive reasoner from scratch is expensive and slow. Copying the traces of a model that already reasons well is cheaper. The gap between those two costs is the entire motive.

This is also why detection is hard and attribution is harder. The users in such a campaign are spread across accounts, regions and payment methods. Tying them to a specific company requires evidence beyond a shared toolset, which is presumably why OpenAI described the link as an association rather than ownership.

The China context

The attribution lands in the middle of a crowded moment for Chinese model releases. Over the National Day holiday, several domestic firms pushed new versions at the same time: Kimi's next model appeared in an API registry, Step 5 Preview opened to third-party endpoints, and Kling put its video model into beta. The market is moving quickly, and the cost of staying near the frontier keeps rising.

Read against that backdrop, the OpenAI disclosure is one signal among several about how the frontier is defended. Labs are treating reasoning extraction as a security problem with a research budget attached, not as a terms-of-service footnote.

There is a second reading, which is that the disclosure itself is a message. Publishing attribution through the Frontier Model Forum and government channels puts other labs on notice about which behaviours will draw a public response.

What the industry does about it

The defences against distillation are mostly technical and mostly unglamorous. Labs can rate limit aggressive querying, fingerprint traffic, and monitor for the patterns that extraction produces. They can also encrypt reasoning traces, which OpenAI did, and then discover that encryption alone does not help if the model can be asked to interpret its own ciphertext.

There is a policy layer too. The Frontier Model Forum exists partly to coordinate exactly this kind of finding between competitors, who have a shared interest in keeping extraction expensive. Sharing through government channels serves a second purpose, which is to give regulators a concrete example of a risk they can act on without writing an entirely new law.

None of this makes distillation impossible, because a model that answers questions can be imitated by a model that reads the answers. It raises the cost, which is the realistic goal. The labs are not trying to end imitation. They are trying to keep it from being cheap enough to skip the training bill.

Why the attribution is the delicate part

Naming a group of people and linking them to a company without accusing the company is a narrow path, and OpenAI walked it deliberately. The wording matters because the alternative readings carry very different consequences. A state-backed campaign would be a diplomatic matter. A group of engineers acting on their own initiative is a corporate governance matter. An ordinary user base pushing at the edges of what the model allows is neither.

The disclosure does not resolve which reading is right. What it does is put the finding on the record in a way that other labs can corroborate, and give regulators a specific incident to point at. That is often the practical purpose of a disclosure like this. It is less a news event than a filing.

There is also a reputational dimension. Distillation complaints have become a routine part of frontier competition, and each one is read differently depending on who is making it. A lab that publishes its evidence, names the pattern and shares it with peers has a stronger position than one that only posts a policy page.

What this does not settle

OpenAI's account does not say what, if anything, was trained on the extracted material, or how much was captured before the pathway closed. It does not name the outside parties beyond the general association. And it does not address whether similar techniques are being used against other providers, which is likely.

Those gaps are not evidence of anything sinister. They are the normal shape of a security disclosure, where the company shares enough to warn the field without publishing a how-to or compromising an investigation.

The part that does travel is the pattern. The most valuable output of a frontier model is no longer just its answers. It is the reasoning underneath them, and the defences around that reasoning are now being tested as deliberately as any system on the network.

Related articles