Claude embeds a text watermark
The mark applies at model level across Claude products, the API, Claude Code, and cloud partners. It may survive copy-and-paste and some editing.
Interactive technical note · ~5 minutes
An LLM can leave a weak, secret-keyed statistical signal across hundreds of completely ordinary token choices. There may be nothing unusual about any one character, word, or token.
00 / THE 30-SECOND VERSION
You do not need to know transformer internals first. The watermark can live entirely in how the next token is sampled.
The model gives several ordinary next-token choices different probabilities.
The key marks some candidates as favored for this context only.
Sampling leans toward that subset without forcing a strange word.
One choice proves nothing. A consistent excess across many choices becomes evidence.
Useful translations: a token is a chunk of text; a logit is a score before probabilities; entropy is how many choices remain plausible; a z-score says how surprising the accumulated excess is.
REAL-WORLD CONTEXT · 11 AUG 2026
Anthropic now says all generated text from supported Claude models carries embedded watermarking, worldwide. It still has not named the method or published enough detail to identify the algorithm.
The mark applies at model level across Claude products, the API, Claude Code, and cloud partners. It may survive copy-and-paste and some editing.
Google DeepMind publishes a production design used in Gemini: secret context-dependent scores steer token selection, and the detector accumulates those scores later.
A detection API is promised, but its statistic and thresholds are not public. The available facts do not identify SynthID, a related method, or an independent design.
Models launched in the EU on or after 2 August 2026 support marking at launch. Earlier models are still being updated.
The model’s mark follows its output across Claude, the API, Code, Cowork, Tag, and supported cloud partners.
Thariq Shihipar says Anthropic will ship a text-detection API that people can use themselves.
Text carries an embedded watermark. Supported files can separately carry signed C2PA provenance metadata.
SAME SPINE · DIFFERENT SAMPLER
The page uses a classic green-list/logit-bias watermark because every step is visible. It is a faithful teaching model for the general idea, not an implementation of SynthID Text.
Primary claims: Anthropic’s marking notice · Thariq Shihipar’s rollout note · SynthID Text in Nature
01 / BLIND TEST
All three came from the same 135M-parameter open model with the same prompt and settings. Two used ordinary sampling; one used the teaching watermark.
The surface text is not the detector.
Trace the revealed passage →There is nothing hidden between the characters.
02 / ONE TOKEN DECISION
The marked passage from 01 is real model output. Here we pause at one position and compare what the model could have said next with what it actually sampled.
The highlight is the token the model actually chose. The table below shows other plausible choices at that exact moment.
Ordering: logits → temperature → green boost → top-k → top-p → softmax → sample.
No token is permanently green.
The class depends on the key and recent context.
Generation sees logits.
The bias is applied before the token is sampled.
Detection sees only text.
It reconstructs the keyed classes afterward.
03 / CHANGE THE REPRESENTATION
Raw text hides the relevant structure. The key turns each eligible token transition into one bit.
In raw text space, nothing jumps out. In the feature space defined by the secret key, the watermark is a biased coin.
04 / THE USEFUL COORDINATE
A generic detector might discover some fingerprint. Here the designer manufactures a specific one—and keeps the key.
05 / ACCUMULATING EVIDENCE
Short answers are not “more human.” They simply contain fewer keyed sampling decisions, so uncertainty is wider.
ONE RUN · READ AS A LEDGER
With a 50% green list, clean text should land near half green. The detector asks whether the observed excess is larger than ordinary random fluctuation.
Not every token is scored: this demo skips the first token and de-duplicates repeated transitions so repetition cannot inflate the evidence.
E[z] ≈ ε√T / √(γ(1−γ))Roughly 4× the text gives 2× the z-score.06 / CAPACITY
When several continuations are plausible, the sampler can steer gently. Near-deterministic decoding leaves little bandwidth without distortion.
The green-logit scheme is pedagogically useful, not uniquely privileged. SynthID Text uses tournament sampling instead: it draws candidate tokens from the model and lets secret-keyed scores decide successive winners. Some configurations can preserve marginal token probabilities more closely—or exactly under idealized assumptions—while encoding detectable correlations.
07 / EDITS + WRONG KEYS
There is no single character to delete. With previous-token context, a local edit usually disrupts only nearby transition scores.
Click any token to replace it. Random replacement is only a mechanical stress test; fluent paraphrasing can be more effective.
08 / SAME PASSAGE, DIFFERENT MAP
The passage is unchanged. Only the representation changed.
Score the passage with key K: does it contain unusually strong evidence for that watermark?
Score the same passage under KA, KB, KC… and look for one convincing match.
The text would carry evidence of a key—not a readable username. “Source” appears only if someone privately maps that key to a source record.
Trying many keys creates more chances for an accidental high score, so a real system needs a stricter threshold. A true multibit watermark instead encodes a payload. Neither extension is implemented here.
09 / THE BROADER IDEA
Neural networks learn transformations that expose useful regularities. This detector uses a keyed transformation designed by hand.
Token IDs are categories, not quantities. ID 4000 is not “twice” ID 2000. Models look up continuous embeddings before computation.
This watermark is an output-distribution relationship. It is not necessarily a latent activation inside the transformer, and it is not SHAP.
INTERACTIVE ANALOGY · NOT THE WATERMARK ALGORITHM
These two groups form an XOR-like relationship. Drag from the raw measurements to a feature that represents their interaction.
The detector is simple once you know the right representation.
WHAT THIS PAGE IS — AND IS NOT
This page implements a canonical teaching watermark. It explains the public claim without pretending to reverse-engineer Claude.