← Back to blog
NewsAbout 7 min read

robots.txt Is a Paper Wall: Reddit v. Anthropic and Australia's Compliance Hearing

Published Oct 7, 2026
robots.txt Is a Paper Wall: Reddit v. Anthropic and Australia's Compliance Hearing

Every website has a small text file called robots.txt. It tells automated programs which pages they may read and which to leave alone. It is a polite sign taped to an unlocked door. The sign has no lock behind it.

That gap between a request and an enforcement mechanism is now the center of two separate fights, one in a US court and one in an Australian parliamentary hearing. Both are asking the same question from different angles: what does a "do not crawl" signal actually mean when it is ignored?

Australia: the opt-out that cannot opt out

Anthropic appeared before Australia's Joint Select Committee on Artificial Intelligence in early October. It had proposed what it called a "narrow form of conditional approval" for AI training on local content, with rights holders opting out by blocking crawlers in robots.txt. The submission implicitly accepted Australia's rejection of a broad text-and-data-mining exemption, and offered the opt-out as the compromise.

The public broadcasters did not accept it. The ABC's head of content and legal operations, Kate Gilchrist, told the committee that the copyright system is adequate and that AI companies can negotiate for the licences they need. Her objection to the opt-out was structural. She said the broadcaster cannot scour the internet to ensure it is opted out on all the relevant sites: "It simply does not work."

SBS made the same point in its own submission, writing that robots.txt signals are routinely bypassed and that the onus should sit with AI companies to avoid wrongdoing.

The asymmetry Gilchrist described is the crux. Rights holders must identify, monitor, and maintain opt-out signals across every possible scraping channel, indefinitely. AI companies face near-zero technical cost to bypass a signal, and until a specific remedy exists, near-zero legal cost as well. A crawler can ignore robots.txt by using a different user-agent identifier, by going through an intermediary, or by having collected the content before the block was placed.

There is a documented basis for the suspicion. TollBit, a startup that brokers licensing deals between publishers and AI companies, found that AI agents from multiple sources were bypassing the robots.txt protocol to retrieve content. Business Insider subsequently reported that OpenAI and Anthropic bypassed robots.txt signals despite public claims of respecting them.

The ABC's own position carries an irony worth noting. In July 2026 the broadcaster selected Anthropic's Claude as its enterprise AI platform, starting a pilot with 100 staff to convert radio broadcasts into digital articles, while simultaneously arguing that Anthropic's products were likely trained on its content without permission. Both things can be true at once, and that is precisely the point: cooperation on one front does not resolve the question on another.

The United States: the contract route

The Reddit case against Anthropic takes a different legal path. A federal court allowed Reddit's contract claims to proceed, which gives platforms a route to challenge data collection outside copyright law.

That distinction matters. Copyright asks whether protected material was copied and whether a defense applies. Contract asks whether the collector accepted conditions around access, automated collection, or commercial use. For platforms whose value sits partly in material they host but do not fully own, the contract theory is the more usable one.

The remand order did not decide that Anthropic breached Reddit's agreement. Its wider importance is narrower: the court treated Reddit's contractual rights as qualitatively different from copyright rights, leaving the breach-of-contract theory alive in the active dispute.

Where terms of service gain practical weight, they turn into data-governance infrastructure. A company needs to prove notice, assent, access conditions, and the relevant version of the terms at the time of collection, beyond what a policy states on its face. Litigation can turn on which terms applied during the disputed access period and how they were displayed.

Why robots.txt was never going to hold

The file has an inverted history. In the 1990s, robots.txt existed mainly to stop search engines from hammering a server too hard. Sites competed to be indexed, because being indexed meant visitors. The relationship has since flipped. The same file is now used to close a door, because the new machines reading the web do not send anyone back.

Blocking works only on crawlers that obey the block. The share of respected news sites blocking AI programs jumped from about 23 percent in September 2023 to nearly 60 percent by May 2025. Yet OpenAI's own crawler was observed ignoring those notes 42 percent of the time in late 2024, and broader testing found violations on 72 percent of UK business sites.

A sign is only as strong as the lawyer standing behind it, and most sites do not have Reddit's.

What this means for builders

For companies building AI systems, the practical lesson is a diligence one. When you buy or collect training data, you need to know whether it was gathered under conditions that conflict with the source platform's terms. That connects contract provenance to the machine-readable trust signals that are supposed to make automatic compliance possible.

The uncomfortable conclusion is that the voluntary protocol is being tested precisely where it was weakest. Reddit built a paywall, sent a legal warning, watched citations of its content rise nearly 40 times in one competitor's answers, and then sued. The sign asked. The gate blocked. Only the lawsuit threatens.

Whether robots.txt grows teeth depends on how the Reddit contract theory fares and on whether Australia legislates a remedy rather than relying on a signal that its own national broadcaster says does not work. Until then, the polite sign on the unlocked door remains exactly that.

The compliance question nobody can answer from the outside

There is a reason the debate keeps circling back to enforcement rather than principle. Everyone agrees in the abstract that a site's access terms should mean something. The disagreement is about who bears the cost of making that true.

Under the opt-out model Anthropic proposed, the cost sits with the rights holder. A publisher must discover which crawlers are reading its content, encode the right blocks, keep them current as new crawlers appear, and monitor for violations. That is an ongoing operational burden that scales with the number of AI companies, and it produces no remedy when a crawler ignores the block.

Under the permission-first model the ABC argued for, the cost sits with the AI company. It must identify what it wants to train on, negotiate for licences, and keep records of what it is allowed to use. That is a familiar arrangement, and it is the one copyright already supports.

Two pale boards standing upright on dark slate with a thin thread of light running between them

The two models put the monitoring cost on different parties, a policy difference with a practical edge, and the party that currently bears it says it cannot afford to. That is why the robots.txt debate keeps resurfacing in every jurisdiction that tries to legislate AI training: the technical standard is being asked to do legal work it was never designed to do.

Why the technical fix keeps falling short

Even a perfect opt-out registry would not solve the underlying problem, because the signal identifies a crawler rather than a company. A firm can scrape through an intermediary, use a different user-agent, or rely on content it collected before any block was placed.

The historical record makes this concrete. GPTBot began formally complying with robots.txt instructions only after the models that made it commercially valuable had already been trained. A prospective block cannot undo a retrospective collection, and a block on one identifier does not bind a firm that can present itself under another.

That is the gap the ABC pointed at when it said it cannot police the entire internet on the rights-holder side. The signal is cheap for the sender and expensive for the receiver, the inverse of how an effective compliance mechanism usually works. Until a remedy attaches to ignoring the signal, the incentive runs one way.

Related articles