Meta Settled the Book Lawsuit. Permission Now Has a Price.

Meta has agreed to settle the class action accusing it of training its Llama models on pirated books. The proposed agreement was filed in federal court in late September, and the notice went out to rights holders days later. The financial terms are sealed until a judge approves them, and Meta is still saying in public that its training practices fall under fair use. What the settlement does not say is as telling as what it does.
The case centered on the acquisition side rather than the training side. Authors alleged that Llama models were trained on works pulled from shadow libraries such as LibGen between 2018 and 2024. That detail matters because a separate ruling already drew a line through the middle of this problem. In the Anthropic case, Judge William Alsup found that training a model on copyrighted books is lawful, but that keeping a library of pirated copies is not. The penalty in that case, a $1.5 billion settlement, attached to how the books were obtained, not to the learning itself.
Read together, the two outcomes describe a rule that is easy to state and hard to live with. Learning from a book you lawfully acquired is fine. Building a corpus out of pirated copies is not. The $1.5 billion figure works out to roughly $3,000 per work, which gives every lab a way to price its own exposure. If you scraped a corpus from a pirate site, you now have a rough number for what that decision costs.
Opt-out is being replaced by a price
The more consequential shift is what happens next. A separate settlement reported this month replaced the voluntary opt-out registry model with a mandatory licensing fee structure tied to how much data gets used. The argument behind it is blunt: opt-out doorways looked fair on paper and failed in practice, because developers rarely checked or honored the exclusion lists. Paying per volume is harder to ignore.
That reframes the entire negotiation. Under opt-out, a rights holder's only lever was to say no and hope. Under a priced license, permission becomes a line item with a number attached, and the argument moves from whether AI companies may use the work to how much they owe for it. For publishers and authors, that is a much better position than a technical barrier nobody enforced.
The practical effect on labs is a new compliance budget. Data provenance stops being an engineering preference and becomes a legal requirement. If one corpus out of a multi-petabyte dataset came from a source that cannot be documented, it introduces a specific liability into the whole training pipeline, and the failure usually happens in the aggregation layer, where third-party datasets are combined without clear attribution.
The output question is still open
Training data is only one of two fronts. The other is what the model produces, and there the picture is less settled. The New York Times survived OpenAI's motion to dismiss by plausibly alleging that ChatGPT's outputs compete with its journalism. That is a market substitution theory, and it cuts differently from the acquisition theory. If the model learns from a work and does not reproduce it, the transformative argument holds. If its output stands in for the original in the market, the argument weakens.
Denying a motion to dismiss is not a finding of liability, and the Times case is still moving. But the two theories together are why AI companies now treat provenance and output monitoring as one problem. A model trained on licensed data can still generate something that competes with the source, and the license terms increasingly say what it may not do.
Audio is not following the same path
If book publishers are heading toward a licensing market, music is heading toward something messier. Suno lost a ruling in Germany brought by GEMA, the performing rights society, over the use of copyrighted recordings and compositions. The American text cases are largely resolving into payouts and retrospective licenses. The audio cases are producing injunctions and jurisdiction fights, which for a product company is the worse outcome, because an injunction stops the service rather than taxing it.
The difference seems to come down to how close the output sits to the input. A language model that summarizes a book is a step removed from the text. A music model that generates a song in an artist's style sits close enough that the comparison is immediate, and courts have been less willing to stretch the transformation argument to cover it.
Where the price does not exist yet
The per-work benchmark is a text phenomenon so far, and the reason is infrastructure. Book publishers have collective licensing bodies, decades of precedent about reproduction rights and a reasonably clear unit of exchange, which is a work. That makes a per-title price easy to define and easy to argue about in court.

Images, video and code do not have the same setup. A stock photo library can price a license per image, but a screenshot of a webpage contains dozens of copyrighted elements at once, and a code repository mixes licensed, permissively licensed and unattributable files in the same tree. Once the question moves from text to any of those, the per-work formula stops applying and the debate returns to fair use, which is the ground AI companies would rather fight on.
The other open question is who collects. A licensing fee that nobody administers turns into a class action anyway, and the settlement structures reported this month still describe payments negotiated case by case rather than a standing market with published rates. Until rates are public, every negotiation happens in private, with the biggest buyers getting the best terms. That is the shape of a market that has formed but not yet matured.
What this means if you build on these models
Most developers reading this do not train foundation models, so the direct exposure is low. The indirect exposure is not. Labs absorbing eight and nine figure settlements pass the cost through, and the mechanism is API pricing. Budget for model access to rise over the next year or two, especially for models trained on licensed material, where the license itself is now part of the cost of goods.
There is a second, quieter implication for anyone hosting or running AI workloads. If you store training datasets for a client, the legal risk now depends on where those files came from, not on what the client does with them. Clear provenance documentation stops being overhead and starts being the thing that keeps an eight figure judgment from landing on a system you operate.
The unsettled part is how far the pricing model spreads. Text publishers have a collective licensing infrastructure and a body of settled cases to point at. Music has societies but a different output relationship. Images, video and code each have their own dynamics, and none of them have a per-work benchmark yet. Meta settling removes one landmark from the map. It does not draw the rest of it, and the next few settlements will decide whether permission becomes a market or stays a series of exceptions.
Related articles
The Memory Chip Is Now the Bottleneck in AI Compute
The logic roadmap is public and predictable. The memory roadmap is gated by packaging yields, a years-long talent base and three companies' capital plans.
AI Music Got a License. Read What It Actually Covers.
Cheap tracks still exist. What is disappearing is the assumption that a track generated from a prompt is free of claims, which was never true and is now demonstrably false.
Hengdian's Studios Are Empty While Asia's AI Film Pipeline Runs
Hengdian's empty stages are the evidence that the industry has already placed its bet, whether or not the audience agrees.
Google Bought 890 Megawatts of Nuclear to Feed Its Data Centers
The compute is available. The electricity is the thing that has to be arranged, years in advance, the way a supply chain is arranged.