← Back to blog
NewsAbout 6 min read

Ninth Circuit Says Copilot's Outputs Are Not Stripped Copies

Published Oct 4, 2026
Ninth Circuit Says Copilot's Outputs Are Not Stripped Copies

A federal appeals court has handed GitHub and OpenAI a narrow but consequential win, and in doing so it drew a line that every company training a model on other people's work should read carefully.

The case is Doe v. GitHub. The plaintiffs are anonymous programmers who published open-source code under licences that require attribution. The defendants are GitHub, its parent Microsoft, and the OpenAI entities. The products at issue are GitHub, Copilot, and OpenAI Codex, all trained on public repositories.

The plaintiffs brought two theories under the Digital Millennium Copyright Act, both built on section 1202, which prohibits removing or altering copyright management information. The first was an input theory: that the defendants stripped attribution data from the plaintiffs' code before feeding it into Copilot as training data. The second was an output theory: that Copilot sometimes returns memorised training data without the source code's attribution, and that doing so amounts to removing or altering that information.

A dim underground archive aisle lined with metal shelves of grey storage boxes

The input theory never got argued

The Ninth Circuit declined to reach the training-stage claim, and the reason is worth noting. In the lower court, the plaintiffs' counsel conceded that copying the training data "perhaps" did not violate the licences, and the district court stated plainly that the complaint "is not about training." The plaintiffs did not object. On appeal, the panel treated the theory as forfeited.

That forfeiture matters more than it looks. The training-stage question is the one that would have touched every model trained on scraped or licensed data. By losing it on a procedural concession rather than on the merits, the plaintiffs left the biggest question open, and left it for a different case.

Why the output theory failed

On the merits, the court rejected the output theory by leaning on the plaintiffs' own description of how Copilot works. The complaint described a "complex probabilistic process" that infers statistical patterns and predicts the most likely completion. That, the panel concluded, describes generation rather than retrieval, and a generated work "cannot reasonably be described as a copy of" the protected code.

The court contrasted Copilot with a search engine. A search engine retrieves stored copies, so it has a copy from which attribution could be removed. A model that predicts the next token does not, in the court's reading, hold such a copy.

The panel then addressed the so-called identicality requirement, and here it did something that will be quoted in briefs for years. It refused to treat identicality as a separate element of a section 1202 claim. Instead, it described identicality as a gloss on the statutory words "remove," "alter," and "copies," and as evidence rather than a test. Near-identical reproduction without attribution supports an inference that attribution was removed. Material differences point the other way, toward a new work to which no attribution was ever attached.

What the ruling actually changes

The practical effect is a narrowing, not a reprieve. Copyright owners who hoped the DMCA would give them a shortcut around the harder questions of fair use and substantial similarity now have to go the long way.

The court's own framing points to where the next fights sit. Owners will lean harder on training-stage theories, arguing that attribution was stripped from data before it ever reached the model. And they will keep pushing ordinary infringement claims about outputs, where the test is similarity to a protected work rather than the mechanics of attribution metadata.

There is a second reading worth holding onto. The panel was careful to say that identicality is evidence, not an element. That gives owners a path: build a record of outputs that reproduce protected code nearly verbatim, attribution absent, and the inference becomes available. The line between a model that has memorised and one that has generalised is now partly a question of how close the output lands, and how well a plaintiff can prove it.

The bigger shape of the problem

Step back and the case looks less like a verdict on AI and more like a snapshot of a legal system catching up unevenly. Attribution was never designed to survive a probabilistic model. Copyright management information assumes a copy with metadata attached. When the system produces something new from a statistical process, there is no copy for the metadata to have been stripped from.

That is an argument about the shape of the technology, and it is the same argument that shows up in every dispute where the old categories do not fit. The Ninth Circuit chose to read the statute's words literally rather than stretch them, which is defensible and also leaves the underlying tension untouched.

For developers, the takeaway is unglamorous. The DMCA is a weaker shield than some had hoped, but it is not gone. For plaintiffs, the takeaway is that the theory you concede in the district court is the theory you lose on appeal. And for everyone training models on public code, the more durable caution comes from the court's own logic: the closer your output lands to a specific input, the less the "new work" framing will hold.

What it means to run a model on someone else's code

For anyone maintaining an open-source project, the practical reading is narrower than the headlines suggest. The DMCA attribution route is harder now, but the underlying licence obligations have not moved. If your code is MIT or Apache licensed, a model that memorised it and reproduces it without the notice is still a licence problem, and the court's own reasoning points at how to prove it.

The discomfort runs the other way too. The ruling relies on a description of Copilot as a probabilistic generator rather than a retrieval system. That description is accurate for most prompts and most outputs. It is less accurate for the tail, where a model reproduces a long, distinctive passage almost exactly, and it is precisely the tail that a plaintiff would need to build a record around. The court has not closed the door on that case. It has made the record the whole argument.

There is also a practical consequence for tooling. The district court's reference to GitHub's own duplicate-detection filter, which flags 150-character matches, is now part of the appellate record. Systems that already measure how close an output lands to training data are the systems best positioned to answer the question the court left open, which suggests the compliance infrastructure and the litigation strategy point the same direction.

The cases still in flight

This decision arrives into a crowded field. Thomson Reuters won its appeal against Ross Intelligence in the Third Circuit in September, on a question of fair use rather than attribution metadata, and that ruling turned heavily on the fact that Ross built a product competing with the very database it trained on. The Ninth Circuit's Copilot decision and the Third Circuit's Westlaw decision are answering different questions, and neither controls the other.

That divergence is the state of AI copyright law in 2026. Courts are resolving the pieces that reach them, in circuit by circuit increments, with facts that do not generalise well. A ruling about a legal research tool says little about a code assistant, and a ruling about attribution metadata says little about fair use. Anyone waiting for a single decision to settle the question is waiting for something that the structure of the system does not produce.

The decision narrows one path and leaves several others open. That is how most of these AI copyright questions are being resolved right now, a piece at a time, in whichever case happens to get there first.

Related articles