← Back to blog
NewsAbout 6 min read

PewDiePie Got Banned Twice for Distilling, and the Rules Just Got Personal

Published Oct 3, 2026
PewDiePie Got Banned Twice for Distilling, and the Rules Just Got Personal

Felix Kjellberg, better known as PewDiePie, told his audience that OpenAI banned him twice. His offense was not smuggling weights or reselling API access at scale. He used outputs from one of OpenAI's models to train Ajax, a 9B model he runs on his own hardware for everyday things like search and email.

Ajax is built on Qwen3.5-9B, a small open model, and fine-tuned on data generated by a larger, closed one. That process is called distillation. It is the same technique that has produced a wave of cheap, capable models over the past two years, and it is the technique the big labs are now working hard to police.

What distillation actually is

Distilling a model means training a smaller one to imitate a larger one. You feed prompts to the teacher, collect its answers, and train the student to reproduce them. The student does not need the teacher's weights or its architecture. It just needs enough of the teacher's behavior to learn the pattern.

This is why the labs care. A frontier model costs tens or hundreds of millions of dollars to train. If a competitor or an individual can harvest its outputs through an API and transfer that capability into a model they own, the expensive part of the work is partly bypassed. The output looks like text, but the underlying signal is the thing being copied.

Many small identical glass bottles being filled from a single overhead glass funnel on a dark surface

Kjellberg's case matters because it is small. He is not a well-funded startup shipping a rival product. He is one person with a consumer GPU and a local model. The enforcement reached that far.

The labs are not guessing about the scale

OpenAI's ban on Kjellberg is the visible tip. The larger campaigns are industrial. Anthropic's chief executive, Dario Amodei, has described a single distillation operation that used around 25,000 fake accounts and ran about 28.8 million exchanges with Claude. At that volume, harvesting outputs is not a side project; it is a data pipeline with staffing and infrastructure.

OpenAI has also moved against companies. It cut off Cursor's access to its models, citing distillation concerns, after xAI used OpenAI outputs to train Grok. The pattern across these actions is that enforcement is no longer aimed only at obvious resellers. It is aimed at anyone whose product depends on capability that was extracted from a closed model.

There is a whole detection industry forming around this. Labs look for account patterns, request volumes and the statistical fingerprints of harvested data. Being caught does not always mean a lawsuit; it can mean losing API access, which for a startup is often the same thing.

Twenty-five companies asked for clearer rules

The labs' enforcement is happening against a backdrop of open disagreement about what the rules should be. In July 2026, twenty-five organizations, including Nvidia, Meta and Microsoft, signed a letter arguing that open-weight models are central to US AI leadership. They asked for targeted legal frameworks that address distillation directly, rather than broad restrictions on open models.

OpenAI, Anthropic and Google did not sign. All three have frontier models whose value depends on limiting how easily their outputs can be copied. The split comes down to who gets to do it and under what terms.

Into that gap, a US appeals court added a separate data point in late September. In Thomson Reuters v. Ross Intelligence, the Third Circuit upheld a ruling that copying a competitor's editorial material to train a search product was not fair use, because the resulting product competed directly with the source. That case is about a search engine rather than a generative model, and the court was careful to say it was not deciding all AI training. But it signals that the legal ground under training data is shifting at the same time the technical enforcement is tightening.

Why the term "distillation" is doing a lot of work

Not every use of a model's output is the same, and the current rules blur the difference. Asking a model to help you write code is normal use. Asking it to generate millions of examples so you can train a rival is not. Somewhere between those two is a person using outputs to make a hobby model behave a little better, which is roughly where Kjellberg sits.

The labs treat output data as an asset they own the rights to use and control. Their terms of service say so, and the enforcement follows the terms rather than a legal standard. That is why the same act can be permitted in one account and banned in another: it depends on volume, on intent as the provider reads it, and on whether the resulting model looks like a competitor. There is no bright line, and the lack of one is the real problem for anyone trying to stay on the right side of it.

Detection makes the boundary harder to see. Labs look for patterns that suggest harvesting: unusual request volumes, accounts created in batches, prompts that resemble evaluation suites. A person running a few thousand prompts may trip none of these. An operation using 25,000 accounts will. In between, the answer is genuinely uncertain, and uncertainty is a poor foundation for a product.

A court decision landed at the same time

The technical enforcement is tightening while the legal picture is still moving. In late September, the Third Circuit upheld a ruling that copying a competitor's editorial material to train a search product was not fair use, because the resulting product competed with the source. The court was careful to note that its decision was about a search engine, not a generative model, and that it was not setting a universal rule for AI training.

Read narrowly, the case says little about Kjellberg's local model. Read as a signal, it says that courts are willing to weigh market harm heavily when a model is trained on material from a product it competes with. The labs are already acting as if that principle applies to their outputs. The courts have not yet said whether it does.

What this means if you run a small model

For individual developers and small teams, the practical picture changed faster than the law did.

Training a local model on outputs you generated yourself sits in a gray zone that depends on the provider's terms of service, not on copyright law. Most major providers prohibit using outputs to train competing models. What counts as "competing" is defined by the provider, and enforcement can arrive without warning in the form of a suspended key.

There are a few things worth doing. Read the terms of service for the specific model you are distilling from, not the general policy. Keep records of how training data was produced, because if a provider asks, an auditable trail is the difference between a warning and a ban. Prefer outputs from models whose licenses explicitly permit derivative training, which includes many open-weight models. And if your product depends on a single provider's API, treat that dependency as a risk, not a convenience.

Kjellberg is an unusual test case because he is public and the stakes for him are low. The lesson is not that he was treated unfairly or that he deserved it. It is that the tolerance for distillation has narrowed to the point where a personal project on consumer hardware drew two bans. Teams running the same technique at scale should assume less margin, not more.

Related articles