← Back to blog
NewsAbout 6 min read

The Watermark Is Not Keeping Up With the Fake

Published Oct 4, 2026
The Watermark Is Not Keeping Up With the Fake

Platforms now label AI-generated content automatically. A new audit suggests the labels are still missing the fakes that matter most.

Researchers audited AI labelling on Instagram, TikTok, X, and YouTube, using the European Commission's July 2026 deepfake guidelines as the reference point. All four platforms applied labels automatically, and a greater share of labels were applied by the platform itself rather than by the uploader, which is progress from earlier audits. The problem is the coverage.

Only 33 per cent of expert-identified deepfakes in systemic risk contexts carried a platform-applied AI label. Those unlabelled items reached a median of 160,000 views, which means a large audience saw them with no warning attached. In controlled uploads of AI-generated content that carried standard provenance signals, the platforms labelled only 61 per cent, and they stripped the hidden signals from 39 per cent of uploads entirely.

That last figure is the one that should worry anyone building on provenance. The label is applied at the surface. The signal that would let someone trace the item later is removed on the way in.

Why the signal disappears

Platforms re-encode media. They resize, compress, strip metadata, and rehost. Every one of those steps is hostile to a fragile watermark, and the systems that attach provenance at generation time cannot control what happens next.

This is the tension the audit exposes. The companies generating synthetic media have largely done the work of signing their output. The platforms distributing it have not done the work of preserving the signature, and in many cases their normal processing removes it before anyone sees the post.

The result is a labelling regime that works when the uploader cooperates and fails when they do not. A user who wants their AI content labelled will usually get it. A user who wants it unlabelled can re-encode the file, and the platform's own pipelines will often help.

What a watermark can and cannot do

Google DeepMind's SynthID is the most widely deployed attempt at solving the attachment problem, and a podcast the lab published on 1 October laid out the design goals with unusual clarity. Hosted by Hannah Fry and featuring Pushmeet Kohli, DeepMind's vice president of science and one of SynthID's architects, along with biosecurity researcher Jeremy Radcliffe, the episode describes what makes a watermark useful: it should be imperceptible to people, hard to remove even after edits and transformations, and easy to integrate and detect at scale.

SynthID's embedder and detector are trained together against an adversary that crops, rotates, shrinks, and adds noise to content, so the mark is built to survive exactly the processing that strips weaker signals. Detection is built into the Gemini app. Partners who share their keys can be checked through a central service. News organisations have used SynthID to verify images circulating online.

That is a strong technical position against accidental removal. It is not a strong position against deliberate removal, and the podcast is honest about the limit. Markers and metadata can be stripped or altered. Text-based markers can be obfuscated by rewriting. A watermark raises the cost of erasure. It does not make erasure impossible.

From images to proteins

The second half of the episode moves somewhere stranger. Radcliffe describes SynthID Bio, a proof of concept that applies the same ideas to AI-designed biological sequences, so that a sequence produced by a model can be told apart from one found in nature.

The purpose is traceability before synthesis. If a design originates from an AI system, marking it means the origin can be established before anything physical is built, which matters for biosecurity in a way that a watermarked photograph does not. The same principle that protects a newsroom checking a circulating image gets extended to a laboratory deciding whether a sequence should exist.

The analogy is not perfect. Protein and DNA sequences have functional constraints that images do not, and the proof of concept is early. But it shows how far the provenance idea can travel once you accept that the question is not "is this real" but "who signed this, and when."

The regulatory clock and the trust gap

The rules have caught up faster than the tooling. The EU AI Act's transparency obligations have applied since 2 August 2026, requiring machine-readable labels for certain AI-generated or manipulated content and audience-visible notices for deepfakes and some material in the public interest. Penalties under Article 50 can reach 15 million euros or 3 per cent of global annual turnover, whichever is higher. Anthropic has announced watermarking for content produced by Claude, joining a list of labs that now mark their output.

Consumer expectations have moved too. In a 2026 Capgemini survey of 12,000 consumers, 67 per cent said they expect companies to clearly indicate when an advertisement shown to them was AI-generated. The demand for labelling is not coming only from regulators.

What the audit suggests is that the label and the trust are drifting apart. Platforms can point to an automatic labelling system and to policy compliance. A viewer looking at an unlabelled deepfake with 160,000 views has no way to know whether the absence of a label means the content is authentic or that the label failed. Absence of evidence is being read as evidence of authenticity, which is the opposite of what a labelling regime is supposed to produce.

The everyday version of the problem

The audit numbers describe systemic risk contexts, but the same machinery governs ordinary content. An AI-generated image of a product, a person, or a place that circulates without a mark is treated as a photograph by almost everyone who sees it, and the correction, if one ever comes, arrives long after the item has travelled.

A separate episode made the point more starkly. An AI-generated image of a nonexistent shoe, attributed to a designer, spread through Hungarian media and social networks and was cited in expert commentary before anyone established that the item did not exist. Nobody was laundering a state secret. They were passing along an image that looked plausible, and the verification step never happened because doing it was harder than sharing.

That is the practical problem watermarking is meant to address, and it is a problem of defaults rather than of determined adversaries. Most people will not strip a mark. They will simply treat its absence as meaningful, whether or not they meant to.

What would actually close the gap

Two changes would move the numbers more than any improvement to watermarking. The first is preservation: platforms would need to carry provenance through their own processing rather than stripping it, which is an infrastructure commitment, not a model problem. The second is a public signal for missing labels, so that an unlabelled item is treated as unverified rather than verified.

C2PA Content Credentials, the provenance standard backed by a membership that spans Adobe, Amazon, the BBC, Google, Meta, Microsoft, OpenAI, and Sony, is the most likely vehicle for the first. It attaches a cryptographically signed manifest at capture or export and lets anyone verify it with open tooling. It cannot detect fakes created before signing, and it does not classify anything as real or fake. What it does is create a chain of custody, which is precisely what gets lost on the way to a feed.

Until that chain survives the upload, the honest summary of AI labelling in 2026 is awkward. The systems work. The coverage does not. And the gap is being filled by whatever the audience decides to assume.

Related articles