Interactive technical note · ~5 minutes

The watermark is
a relationship.

An LLM can leave a weak, secret-keyed statistical signal across hundreds of completely ordinary token choices. There may be nothing unusual about any one character, word, or token.

Start with the blind test
keyed score z = — unlikely under the clean null
01RAW TEXTordinary characters
02TOKENScategorical IDs
03φK(context, token)secret-keyed feature
04COUNTone-dimensional score

00 / THE 30-SECOND VERSION

Four ordinary steps. One unusual relationship.

You do not need to know transformer internals first. The watermark can live entirely in how the next token is sampled.

  1. 1Predict

    The model gives several ordinary next-token choices different probabilities.

  2. 2Label privately

    The key marks some candidates as favored for this context only.

  3. 3Nudge

    Sampling leans toward that subset without forcing a strange word.

  4. 4Count later

    One choice proves nothing. A consistent excess across many choices becomes evidence.

Useful translations: a token is a chunk of text; a logit is a score before probabilities; entropy is how many choices remain plausible; a z-score says how surprising the accumulated excess is.

REAL-WORLD CONTEXT · 11 AUG 2026

Claude is marking text. The recipe is still private.

Anthropic now says all generated text from supported Claude models carries embedded watermarking, worldwide. It still has not named the method or published enough detail to identify the algorithm.

CONFIRMED

Claude embeds a text watermark

The mark applies at model level across Claude products, the API, Claude Code, and cloud partners. It may survive copy-and-paste and some editing.

DOCUMENTED EXAMPLE

SynthID Text is real and deployed

Google DeepMind publishes a production design used in Gemini: secret context-dependent scores steer token selection, and the detector accumulates those scores later.

STILL UNKNOWN

Claude’s exact sampler

A detection API is promised, but its statistic and thresholds are not public. The available facts do not identify SynthID, a related method, or an independent design.

ROLLOUT New models first

Models launched in the EU on or after 2 August 2026 support marking at launch. Earlier models are still being updated.

REACH Worldwide once supported

The model’s mark follows its output across Claude, the API, Code, Cowork, Tag, and supported cloud partners.

VERIFICATION A detection API is coming

Thariq Shihipar says Anthropic will ship a text-detection API that people can use themselves.

TWO DIFFERENT MARKS Text signal ≠ file metadata

Text carries an embedded watermark. Supported files can separately carry signed C2PA provenance metadata.

SAME SPINE · DIFFERENT SAMPLER

Change the implementation, keep the intuition.

1CONTEXT + KEYmake a repeatable secret signal
2FAVOR A KEYED SUBSETadd a small logit boost, then sample
3ORDINARY TEXTno hidden characters required
4COUNT FAVORED TOKENScompare the excess with a null model

The page uses a classic green-list/logit-bias watermark because every step is visible. It is a faithful teaching model for the general idea, not an implementation of SynthID Text.

Primary claims: Anthropic’s marking notice · Thariq Shihipar’s rollout note · SynthID Text in Nature

01 / BLIND TEST

Which passage carries the watermark?

All three came from the same 135M-parameter open model with the same prompt and settings. Two used ordinary sampling; one used the teaching watermark.

PASSAGE A

PASSAGE B

PASSAGE C

1 marked · blind chance = 1 in 3
THE POINTOne guess cannot establish visual detection.

The surface text is not the detector.

Trace the revealed passage →

02 / ONE TOKEN DECISION

Now zoom into one decision inside it.

The marked passage from 01 is real model output. Here we pause at one position and compare what the model could have said next with what it actually sampled.

secret key+previous token deterministic predicate ◆ favored / ○ not favored
MARKED PASSAGE FROM 01 · TOKEN —

The highlight is the token the model actually chose. The table below shows other plausible choices at that exact moment.
candidateclassprobability

Ordering: logits → temperature → green boost → top-k → top-p → softmax → sample.

No token is permanently green.
The class depends on the key and recent context.

Generation sees logits.
The bias is applied before the token is sampled.

Detection sees only text.
It reconstructs the keyed classes afterward.

03 / CHANGE THE REPRESENTATION

Reveal the feature.

Raw text hides the relevant structure. The key turns each eligible token transition into one bit.

In raw text space, nothing jumps out. In the feature space defined by the secret key, the watermark is a biased coin.

04 / THE USEFUL COORDINATE

Obvious statistics mostly overlap.

A generic detector might discover some fingerprint. Here the designer manufactures a specific one—and keeps the key.

Clean and watermarked corpus comparison Ordinary aggregate features overlap, while keyed z-scores separate.
cleanwatermarkedEntropy and top-token share overlap heavily.

05 / ACCUMULATING EVIDENCE

One token tells you nothing. Five hundred weak choices can.

Short answers are not “more human.” They simply contain fewer keyed sampling decisions, so uncertainty is wider.

ONE RUN · READ AS A LEDGER

Count the excess, not any special token.

With a 50% green list, clean text should land near half green. The detector asks whether the observed excess is larger than ordinary random fluctuation.

SCORED CHOICES—
CLEAN GREEN RATE50%
EXPECTED GREEN—
OBSERVED—
EXPECTED—
EXTRA GREEN VOTES—
STANDARDIZED—

Not every token is scored: this demo skips the first token and de-duplicates repeated transitions so repetition cannot inflate the evidence.

Across many generated passages
cleanmarkeddashed line: example threshold z = 3
The same run, token by tokenevidence can rise or fall locally
Signal adds with more choices; random noise grows more slowly.E[z] ≈ ε√T / √(γ(1−γ))Roughly 4× the text gives 2× the z-score.

06 / CAPACITY

The signal lives in the model’s freedom to choose.

When several continuations are plausible, the sampler can steer gently. Near-deterministic decoding leaves little bandwidth without distortion.

entropy — green mass — → — KL / token —

The green-logit scheme is pedagogically useful, not uniquely privileged. SynthID Text uses tournament sampling instead: it draws candidate tokens from the model and lets secret-keyed scores decide successive winners. Some configurations can preserve marginal token probabilities more closely—or exactly under idealized assumptions—while encoding detectable correlations.

07 / EDITS + WRONG KEYS

Distributed evidence degrades in pieces.

There is no single character to delete. With previous-token context, a local edit usually disrupts only nearby transition scores.

ORIGINAL—
AFTER EDITS—

Click any token to replace it. Random replacement is only a mechanical stress test; fluent paraphrasing can be more effective.

08 / SAME PASSAGE, DIFFERENT MAP

Try the wrong key.

The passage is unchanged. Only the representation changed.

CORRECT KEY—
DETECTOR KEY—
OPTIONAL EXTENSIONCould the same signal identify a source key?
Possible in principle; not part of the main demo. This is not a claim that Claude embeds a user, account, request, or generation ID.
ONE SHARED KEY

Detection asks a yes/no question.

Score the passage with key K: does it contain unusually strong evidence for that watermark?

passage+key K→marked?
DIFFERENT KEYS FOR DIFFERENT SOURCES

Attribution would search candidates.

Score the same passage under KA, KB, KC… and look for one convincing match.

KAweak KBmatch KCweak
matching key KB→provider’s private lookup table→source record B

The text would carry evidence of a key—not a readable username. “Source” appears only if someone privately maps that key to a source record.

Trying many keys creates more chances for an accidental high score, so a real system needs a stricter threshold. A true multibit watermark instead encodes a payload. Neither extension is implemented here.

09 / THE BROADER IDEA

A hand-designed feature map.

Neural networks learn transformations that expose useful regularities. This detector uses a keyed transformation designed by hand.

RAW COORDINATESpixels · token IDs · surface countsstructure can be obscure
FEATURE SPACEedges · concepts · keyed token relationsthe useful regularity is exposed
DECISIONline · probe · one scoreeasy after the transformation

Token IDs are categories, not quantities. ID 4000 is not “twice” ID 2000. Models look up continuous embeddings before computation.

This watermark is an output-distribution relationship. It is not necessarily a latent activation inside the transformer, and it is not SHAP.

INTERACTIVE ANALOGY · NOT THE WATERMARK ALGORITHM

The pattern can become a coordinate.

These two groups form an XOR-like relationship. Drag from the raw measurements to a feature that represents their interaction.

Nonlinear feature-map analogy Two groups that are tangled in raw coordinates separate after their interaction becomes a feature.
relation holdsrelation does not holdNo straight boundary separates the groups in the raw x₁/x₂ view.
The detector is simple once you know the right representation.
ADVANCED MODERun the browser simulator

clean mean z—
watermarked mean z—
TPR at z ≥ 3—
FPR at z ≥ 3—

WHAT THIS PAGE IS — AND IS NOT

What this accurately explains

  • Generation-time token selection can carry a statistical watermark.
  • Small key-dependent biases can accumulate into detectable evidence.
  • Detection can use only text, the key, and the algorithm.
  • Longer, higher-entropy outputs usually provide more evidence.

What remains unknown about Claude

  • Anthropic has not yet disclosed the exact embedding method or detector internals.
  • A user-facing detection API is promised but not yet documented.
  • There is no confirmation that Claude uses SynthID Text.
  • Claude may not use green lists or a constant logit boost.
  • There is no public evidence of per-user or per-request payloads.
  • This demo cannot detect actual Claude output.

This page implements a canonical teaching watermark. It explains the public claim without pretending to reverse-engineer Claude.