← Back to blog
NewsAbout 6 min read

Thomson Reuters v. Ross: The First Appeals Ruling on AI Training and Fair Use

Published Oct 2, 2026
Thomson Reuters v. Ross: The First Appeals Ruling on AI Training and Fair Use

For three years, the legal fight over whether AI models can be trained on copyrighted material has mostly either settled or stalled at the district level. In September, the Third US Circuit Court of Appeals changed that. Its decision in Thomson Reuters v. Ross Intelligence is the first federal appellate ruling holding that using copyrighted material to train a competing commercial AI tool does not qualify as fair use.

Judge Tamika Montgomery-Reeves affirmed the lower court's February 2025 finding that Ross Intelligence infringed 2,243 Thomson Reuters headnotes, which are short editor-written summaries of legal points, when it trained a competing legal search product on them. The appellate stamp matters because it lifts the reasoning from one judge's opinion into binding precedent for a federal circuit, which is the kind of thing other courts now have to reckon with.

What the case actually decided

The narrow holding is about a specific kind of system. Ross built a legal research tool intended to compete directly with the product whose content it trained on. The court found that this use failed the fair-use test, largely because the output competed in the same market as the original.

The reasoning leans on market substitution. When the purpose of training is to produce something that replaces the source, the fourth fair-use factor, the effect on the potential market for the original, weighs heavily against the user. The court was not persuaded that the intermediate copying involved in training was transformative enough to overcome that.

It is worth being precise about what this does and does not cover. Ross was not, in the relevant sense, a generative model producing original expression. It was a search and retrieval product that reproduced and repurposed copyrighted summaries. The court drew that distinction explicitly, leaving room for generative AI to argue for broader fair-use protection on the grounds that it generates new content rather than reproducing existing content.

That carve-out is generous on its face and less comforting in practice. It means the appellate ruling does not settle the question the image and video industry most wants answered: whether training a diffusion model on billions of copyrighted images and videos is fair use. It leaves that question to future cases, and it means the answer may differ by model type.

The generative-AI question is still open

The parallel fight on the generative side has been moving through different courts, and it has not been resolved in the industry's favor. The most discussed data point is the settlement in Bartz v. Anthropic, in June 2025, which resolved a class claim over training data for a reported $1.5 billion. A settlement is not a ruling, but the size of the number tells you how the defendant's lawyers assessed their odds.

Alongside the appeals ruling, a separate case over AI image training went to a jury trial in San Francisco, making it the first such matter to reach a jury rather than settling. A verdict there would give the industry something no party has had yet: a factual determination about whether training image models on artists' work infringes, decided by people rather than a judge.

Read together, the Ross ruling and the ongoing cases point in one direction. Courts are increasingly unwilling to treat "we trained a model on it" as an automatic defense, and they are separating the analysis by how closely the resulting product competes with the source. The more directly a system substitutes for the content it learned from, the weaker the fair-use argument becomes.

Why this reaches beyond legal tech

It is tempting to file Ross under a narrow category, legal research, and move on. That would be a mistake for anyone using AI to produce commercial content, and the reason is the market-substitution logic.

Consider the ordinary case of an AI tool that writes product descriptions. If it was trained on a retailer's catalog copy and then used to generate competing catalog copy, the Ross reasoning applies with uncomfortable directness. The same holds for an image model trained on the product photos of the brand it is helping a competitor undercut, or a copywriting tool fed on a publication's articles and used to generate content that draws the same readers.

The practical exposure is not limited to model developers. Businesses that deploy AI tools inherit risk from how those tools were built, and the audit trail is often missing. Sellers working with free models trained on unlicensed data are the most exposed, because they cannot demonstrate provenance even if they wanted to. Teams that source models trained on licensed or owned data can document that choice, which is a defensible position in a dispute and increasingly a requirement imposed by partners and platforms.

The compliance cost of this shift is real and it is uneven. Larger players can negotiate licenses and build licensed datasets; smaller ones cannot, and face a choice between paying for licensed tools, generating original content by hand, or assuming a risk they may not be able to quantify. Platforms are likely to fill part of that gap themselves, scanning uploaded content for signals of unlicensed AI-generated material the way they already scan for counterfeits.

The practical ambiguity is worth stating plainly. The Ross ruling covers a retrieval product, not a diffusion model, so nobody should read it as a verdict on image or video training. What it does establish is a method of analysis that will be applied more broadly: courts asking how closely the trained system's output competes with the material it learned from. That question has an uncomfortable answer for a lot of commercial AI products, and it does not depend on whether the model is generative.

The near-term consequence is that provenance is turning into a product feature. Tools that can show where their training data came from, or that were built only on content the customer already owns, carry less risk and can charge for that. Tools that cannot will drift toward customers who either do not know to ask or cannot afford to care.

The state of play

What exists now is a set of markers rather than a settled rule. Training on copyrighted material to build a directly competing commercial product is not fair use, at least in the Third Circuit. Generative models have more room to argue, and that room has not been tested to a jury verdict. Settlements have priced the risk at figures that get attention in boardrooms.

For anyone building products on top of AI, the sensible response is to stop waiting for the law to crystallize. Know where the training data came from, keep that record, prefer models that can document their provenance, and treat the source of a model's weights as a supply-chain question rather than a legal footnote. The direction of travel is clear even if the destination is not, and the organizations that can answer the provenance question today will spend far less energy answering it under a deadline later.

Related articles