OpenAI Will Watermark ChatGPT Text in the EU. Its Own Numbers Show How Fragile That Is

On October 5, OpenAI said it will add an invisible watermark to text produced by ChatGPT and Codex in the European Union, rolling out over the coming weeks to eligible users on all plans. The method, called textGrain, does not insert a visible symbol or a hidden character you could delete with find and replace. It subtly shapes the model's next-word choices using a secret key, stacking hundreds of statistical nudges that a detector can later recover from the text alone. Because the pattern lives in the wording, it survives copy and paste.
The trigger is regulation. Article 50 of the EU AI Act requires providers of generative systems to mark their output in a machine-readable form. Those transparency obligations took effect on August 2, 2026, and the European Commission has confirmed that systems already on the market before that date have until December 2, 2026. The Commission published a Code of Practice on marking and labelling AI-generated content in June, and roughly 190 organizations had signed by the end of July. Anthropic said in August it would watermark Claude text worldwide; Google has SynthID for text. OpenAI is arriving on time for a deadline it spent two years avoiding.
The company published its own failure curve
The most useful part of the announcement is the list of numbers OpenAI put next to it. At a target false-positive rate of 1%, on English text drawn from the ELI5 dataset, the detector identified the watermark in about 80% of 200-token passages and about 95% of 400-token passages. Mathematics scored substantially lower, because the model has less freedom in choosing words. Replace 10% of the words with synonyms in a 400-token passage, and detection falls from roughly 92% to 66%. Replace 25%, and it drops to about 17%.
Read those rows in order and the picture is clear. The detector works best on long, loosely constrained prose. A short email is closer to 200 tokens than 400. Code and math leave little room to hide a signal. A light pass through any paraphrasing tool pushes detection toward a coin flip.

OpenAI lists five limits, and they are worth quoting in any policy document you write this quarter. The watermark does not measure human contribution. It does not establish ownership or responsibility. It does not identify the user, since no account, prompt, or conversation is tied to the text. It does not verify accuracy, and a watermarked passage can be false. Finally, the absence of a watermark does not prove a human wrote the text, because the passage may be too short, might have been edited or translated, could come from an unsupported model or from before the rollout, or might simply come from a different company's tool.
That last point is the one that will cause real damage in practice. Someone will drop a student's essay or a job candidate's cover letter into a detector, get no match, and treat the result as proof of honest work. Or get a match and treat it as proof of cheating. Neither conclusion follows, and OpenAI says as much in its own post.
Why the detector is not public
Access to the detection tool is limited at launch to approved researchers and expert organizations, granted case by case. That caution is defensible, since a tool with a 66% detection rate after a synonym swap is easy to misuse. It also means independent verification is hard while the system is being rolled out, which is an awkward place for a compliance feature to start.
There is history here. In August 2024, OpenAI declined to watermark ChatGPT after an internal survey found that nearly 30% of users would reduce their usage if it did. The current arrangement is a compromise: EU-only at first, off by default for API customers everywhere, who can opt in for select models. Globally, text watermarking is not the default.
What this changes for anyone who ships text
For teams in the EU, the practical change is smaller than the headline suggests. Staff output will carry a signal no one can see. It does not identify them, does not degrade model quality in OpenAI's measurements, and follows the text into any document or client email it is pasted into. If someone with detector access later inspects that text, they may be able to flag it. Nothing about the workflow needs to change.
The real risk sits further down the chain, in any process that treats provenance detection as a verdict. Hiring, admissions, newsroom verification, and academic integrity systems are the obvious candidates, and each of them contains a human who will be tempted to read a detection result as evidence. OpenAI's own numbers argue against that. A marker that routine paraphrasing can strip may satisfy the letter of Article 50 while failing the job the article was written to do.
Enforcement gives the deadline teeth: a violation can cost up to 15 million euros or 3% of global annual revenue, whichever is higher. That is why the rollout is happening. But the more interesting question is whether a watermark that fades under light editing can carry the weight regulators are placing on it. The December 2 deadline for legacy systems, the eventual public release of the detector, and any Commission guidance on acceptable detection thresholds are the milestones to watch.
For now, the safest rule for any team is the one OpenAI wrote into its own post: a missing watermark does not prove a human authored a text, and a present one does not prove the model wrote all of it. Watermarks can indicate that a system touched a passage. They cannot tell you how much of the thinking was a person's.
How text watermarking actually works
It helps to see why the signal is so easy to lose. A conventional watermark is a mark attached to a file, and deleting it is a matter of finding the right field. textGrain takes a different route. During generation, a secret key biases the model's choice among near-equivalent next words, nudging it toward a subset that a detector holding the same key can recognize. Across hundreds of such choices, the pattern becomes statistically visible. Because the bias lives in the wording, the mark travels with the text through copy and paste, through a different editor, and through a different app.
The fragility follows from the same design. Any process that rewrites the words, whether a paraphrasing tool, a translation, or a heavy human edit, disturbs the pattern the detector looks for. That is the trade at the heart of every text watermark: a mark strong enough to survive light editing is also visibly different from ordinary writing, which defeats the point. OpenAI chose a subtle mark and published the resulting detection curve, which is more honest than most vendors would be.
The comparison across labs is instructive. Anthropic said in August it would watermark Claude output worldwide, and its move drew pushback from users who argued that instructions, context, and decisions come from the person, with the model acting as a tool. Google's SynthID does something similar for its own text. The three approaches amount to an industry convergence on the same idea, with the same limits, which suggests the technique is close to the ceiling of what is currently possible for text.
Meanwhile, the harder problem sits untouched. Marking text is comparatively easy because a passage contains hundreds of word choices to encode a signal into. Images and audio carry the same kind of statistical room, and OpenAI already offers verification for them. But none of these markers answers the question a newsroom or a school actually needs answered, which is whether a specific claim is true and who stands behind it. Provenance and accuracy are separate problems, and watermarking addresses only the first.
Related articles
The Money Is Moving Into Physical AI: SiMa.ai's $1.45B, and Agents That Design Hardware
The capital is arriving ahead of the evidence. The deployments over the next year will say whether the bet was right.
An Arizona Court Threw Out a Sentence Because an AI Video Spoke for the Victim
Reproducing what a person did is one thing. Speaking for them is another, and the line between the two is now written into a court record.
MCP's Trust Model Let One Poisoned Agent Reach Into Another
The instruction never crosses a boundary in the code, so no layer is left to ask whether the request made sense.
AI Anime Short Dramas Obtain the First Data Intellectual Property Certificate: How the Creative Process Becomes Evidence for Rights Protection
When the marginal cost of content laundering approaches zero, what creators lack most is not ideas, but a record that can prove where the ideas came from.