← Back to blog
NewsAbout 6 min read

Anthropic Will Now Bill You for a Request It Refuses to Answer

Published Oct 4, 2026
Anthropic Will Now Bill You for a Request It Refuses to Answer

On 24 September, at 17:11 UTC, Anthropic's developer account posted two paragraphs that changed a line in the billing documentation. The company will resume charging for requests its safeguards block before Claude responds. The change applies only in three categories where it measures low false-positive rates: biology, frontier large-model development, and reasoning extraction, which the post calls distillation attacks.

The mechanics are simple and slightly uncomfortable. When a Claude classifier declines a request before generating anything, the API does not return an error. It returns a normal HTTP 200 with `stop_reason: "refusal"`, an empty content array, and a `stop_details` object whose `category` field names the policy area. The refusal carries no content, but token counts still appear in usage, and the request still counts against your rate limits. Under the new rule, in those three categories, the refusal is also charged, at the rates of the model that ran it.

Anthropic had removed charges for zero-output refusals on 2 June, around the launch of its Fable product. Mid-stream refusals, where the model starts answering and is then stopped, were always billed. Refusals before any output in other categories, including cybersecurity and general harm, remain free. The change applies across the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry, and the documentation adds a warning that the billed categories may change again.

Why billing is now a security control

The company's stated reason is that it has seen coordinated attacks on its systems in recent weeks. A blocked request that costs nothing is a free probe. An attacker who wants to map the shape of a classifier can send thousands of requests and learn from the refusals without paying a cent. Putting a price on the blocked attempt in exactly the three categories under attack raises the cost of that mapping, and Anthropic is explicit that the charge is meant to work as a layer of defence rather than as revenue.

Two numbers came with the announcement. In recent testing, 99.7 percent of accounts using Claude Code, Claude.ai or Cowork did not hit any of the newly billable blocks. The classifiers behind them are tuned to a false-positive rate below 0.1 percent, and the post adds that the figure is not zero and that the classifiers will keep improving so they interrupt work less often.

The classifiers run on four Claude models: Fable 5.1, Fable 5, Opus 5.5 and Opus 5. The documentation names five categories. Two of them, cybersecurity and general harm, are not billed. The three that are billed map onto work that Anthropic's commercial terms already restrict, which is the reason the false-positive rate is low enough to measure in the first place: few legitimate customers are running frontier model training or asking a model to reproduce its own internal reasoning.

A clear glass vial with a black cap standing on a dark reflective surface

Distillation attacks are worth spelling out, because the name is doing a lot of work. The category code is `reasoning_extraction`, and the behaviour it targets is prompting a model to expose step-by-step internal reasoning so that the output can be used to train a competing system. It is the mechanism behind several of the model-theft disputes of the past year, and it is difficult to separate from ordinary users who simply want the model to think out loud.

The false positives are already visible

Open issues on the claude-code GitHub repository tell a different story about the rate. One user logged 23 flags in the reasoning-extraction category on ordinary work, including podcast copy and budget spreadsheets, and said seven of them terminated the turn. Other reports describe standard software development tasks and Claude Code's own workflow documentation triggering the same block. Anthropic has not explained how a refund would work for a block that turns out to be wrong, beyond pointing users at the `/feedback` command.

That gap is the practical problem. A developer can count `stop_reason: "refusal"` events and retry against a fallback model, which is the documented workaround, and the fallback credit arrangement is unchanged. What they cannot easily do is tell, at the moment of the charge, whether the block was correct. If the answer arrives later, the money has already moved, and a refund policy that has not been published is not something a finance team can budget against.

The response to the post was moderate rather than angry. In roughly three hours it collected about 1,511 likes, 64 reposts and 220 replies on X, and the Hacker News submission drew five points and two comments. Most of the discussion was about the principle rather than the amount, which suggests the amount is small enough that few teams will notice it on an invoice.

What a team can actually control

Two things in the documentation are worth wiring into a pipeline. The first is the refusal category, which arrives as a machine-readable field, so a service can separate refusals from errors in its logs and count them per category over time rather than discovering the pattern during a quarterly review. The second is the fallback path. The documented pattern is to retry a refused request against a fallback model, and Anthropic says the fallback credit arrangement is unchanged, which means a retry does not pay twice for the same piece of work.

Neither of those solves the false-positive problem. When a legitimate request is blocked in a billed category, the record shows a refusal with a category code and no content, and the invoice shows a charge. A team that never reads that field will find out the rate only when somebody asks why the bill moved.

The principle is the interesting part

Every large model provider runs some version of this calculation. Refusals consume compute, and a refusal that a customer can trigger at will is a lever an attacker can pull. Charging for the refused request moves the cost back onto whoever is pulling it.

The trouble is that the same lever is reachable by people who are not attacking anything. Someone writing a biology paper, or a startup working on model distillation techniques, sits in the billed categories by definition of the work rather than by intent. Anthropic's own framing is that the false-positive rate is low enough to accept, which is a judgement made by the party that collects the fee.

There is a cleaner design available, and Anthropic has partly built it. The `stop_details.category` field is machine-readable, so a team can route refusals into a log, count them by category, and audit them in aggregate. Refused responses also appear in usage records with token counts attached, which means the data a dispute would need is already being recorded. Whether that becomes a normal part of running Claude in production, or stays a footnote most teams never read, decides how much the billing change costs above the invoice line.

Related articles