Anthropic said this week it will embed an imperceptible machine-readable watermark in text generated by Claude, and attach cryptographically signed provenance metadata to generated files.
The requirement attaches to models launched in the European Union after 2 August, though the company says the marks will travel with output wherever Claude is offered.
This is Article 50 of the EU AI Act doing what it was designed to do, which is to make provenance a product feature rather than a research topic.
Whether it will work depends entirely on what you think the problem is.
What the law actually asks for
Article 50 became enforceable on 2 August and applies regardless of when a system reached the market, with penalties reaching €15 million or 3% of worldwide turnover.
Providers of systems that generate synthetic text, audio, image or video must mark outputs in machine-readable form and enable their detection.
Systems already on the market before that date have until 2 December to comply, which is where the deadline in Anthropic's announcement comes from.
There are carve-outs for assistive editing and for cases where the system does not substantially alter the input or its meaning, and the Commission's voluntary Code of Practice on Transparency offers signatories a presumption of conformity.
None of this obliges anyone to catch bad actors.
The obligation is to mark and to make marks detectable, which is a narrower thing than making AI-generated content identifiable in the wild.
The honest limits, stated honestly
To its credit, Anthropic has published the caveats rather than burying them: heavy rewriting, mixing Claude's output with other copy, or working in very short passages can defeat detection, and AI-assisted edits to human text can register as AI-generated.
The academic literature is blunter still.
Google's SynthID-Text, the only comparable system deployed at scale, has been shown to degrade substantially under paraphrasing, copy-and-paste modification, back-translation and synonym substitution.
Researchers at ETH Zurich found that off-the-shelf paraphrasers scrub it more easily than rival schemes, and that black-box queries to steal the watermark first push removal towards complete success.
A recent survey of the field summarises text watermarking as fragile to paraphrase, with a capacity of roughly one bit.
The theoretical position is worse: it has been proved that no watermark is secure against a determined white-box adversary.
So the mark catches the student who pastes an essay unaltered, and misses the disinformation operation that runs everything through a second model.
The metadata half fares no better
Signed file provenance, essentially the C2PA approach, has an even more familiar failure mode.
Every major social platform recompresses uploads, and in doing so discards embedded metadata as a matter of routine engineering rather than malice.
One test of six platforms found C2PA manifests fully stripped in five, and Meta has been removing EXIF data from photographs for well over a decade.
The paradox is exact: the content that most needs provenance, because it is spreading virally, passes through precisely the pipelines that destroy it.
LinkedIn preserves credentials, Cloudflare Images survives its own transformations, TikTok manages partial preservation. That is not an ecosystem, it is a handful of exceptions.
Durable Content Credentials, which pair the manifest with an invisible watermark and a perceptual fingerprint, are the industry's answer, and they mitigate rather than solve.
So what is it for
The temptation is to conclude that this is theatre. It is not, but it is worth being precise about what it buys.
Watermarking is a floor, not a filter.
It makes casual undisclosed use detectable, gives platforms a signal they can act on cheaply, and hands compliance teams something auditable.
It creates an asymmetry that matters commercially: enterprises can now demonstrate what came from a model, which is useful in regulated sectors and in litigation.
And it establishes a norm that provenance is the vendor's responsibility rather than the reader's problem.
What it does not do is address the slop.
The content flooding Reddit, Spotlight feeds and search results is not generated by people who will leave a watermark intact.
It is produced at scale by operators with every incentive to paraphrase, chunk and launder, and the tools for doing so cost nothing.
Meanwhile 2 December approaches, and the more revealing question is not whether Anthropic's marks hold up.
It is which of its competitors ship anything at all.