How the invisible watermark inside Claude actually works
No hidden characters, no extra words: the marking lives in the word choices themselves. Anthropic has now described the method in detail.

Illustration · AI-generated (AI IN LIFE)
At a glance
- Method: statistical shift in word choice, no hidden characters, no additional tokens
- Basis: SynthID-Text, an open-source method from Google
- Files additionally carry signed provenance data in the C2PA standard
- The marking survives copying and light editing
- Trigger: EU AI Act labelling obligation since August 2026; rollout is global and across all products
Anthropic has explained how text watermarking in Claude works. The method inserts no invisible special characters and generates no additional tokens. Instead, during generation the model shifts its word choices very slightly along a predetermined pattern. At any single point this is meaningless; across enough text it produces a statistically detectable signature.
The technical basis is SynthID-Text, an open-source method from Google. For files such as images, signed provenance data in the C2PA standard is added on top. According to Anthropic the text marking survives copying and pasting as well as light editorial changes — it travels with the text when the text is reused.
The trigger is the labelling obligation of the European AI Act, in force since August, along with its associated code of practice. Anthropic is applying the marking not only in the EU, however, but globally and across the product line — from programming interface access through to end-user applications.
What matters in practice are the limits the company itself names. A detected watermark shows that Claude was involved in a text — not that Claude conceived it. Models are frequently used to correct, shorten or translate other people's writing. The detection therefore says something about the tool and little about authorship.
Equally important is what an absent watermark means: nothing. Anyone using a provider without marking, running a model themselves, or heavily rewriting the output leaves no signature. A test result can incriminate but can never exonerate — an asymmetry that leads to false conclusions quickly in schools and newsrooms.
That leaves the practical use. For platforms that need to sort provenance at scale, a statistical marker is far more robust than any guessing detector. For judging a single text it is poorly suited. The method solves an infrastructure problem, not a trust problem.
We explained how the method works technically in an earlier piece.
FAQ
Does the watermark visibly change the text?
No. No characters are inserted and no extra words generated. Only the probability of choosing particular phrasings is shifted.
Does a hit prove an AI wrote the text?
No. It shows Claude was involved — correcting, shortening or translating a human text also produces the signature.
What does it mean if no watermark is found?
Nothing reliable. Other providers, self-hosted models or heavy rewriting leave no signature. Absence is not exoneration.


