On August 2, 2026, Article 50 of the EU AI Act took effect: providers of generative AI systems must ensure their outputs “are marked in a machine-readable format and detectable as artificially generated or manipulated” (Article 50). On August 11, Anthropic announced it would embed invisible watermarks in Claude’s text output. A technical note followed on August 14 with the design: the scheme is Google DeepMind’s SynthID-Text; it applies worldwide, because Anthropic does “not yet have a durable way to scope it by region”; it is on by default, and the note mentions no way to turn it off. Google got there earlier: it announced SynthID text watermarking for the Gemini app and web in May 2024, and the production integration was documented that October in Nature.

The first days after launch produced a wave of “we’ll get caught” panic, which I took apart in my earlier watermark post. My conclusion then: the watermark doesn’t filter for the people most determined to cheat, it filters for the people least on guard, and the boundary of what counts as a violation is for schools and employers to fill in. A month later the argument hasn’t cooled, but its center has moved. A paper posted to arXiv on September 9 (arXiv:2609.09604, by Alexander Nemecek and colleagues) gathers the representative positions from a month of public fighting; the authors flag that this is an illustrative reconstruction rather than a systematic survey, with no claim about which view is most common. Some users say the watermark ruins output quality, code above all. Others say it hides information that can be traced back to individuals. The paper records a GitHub tool claiming to strip the watermark collecting more than ten thousand stars within a week, and business-press reports of users canceling subscriptions in protest. Then it points at the awkward fact: these complaints contradict each other. A watermark that comes off easily can hardly be an inescapable tracker at the same time.

The paper’s real point sits elsewhere, though: neither camp can back up its claims. Anthropic says internal testing showed “no impact of watermarking on the content, level of creativity, or readability of Claude’s text,” and that the watermark “carries no identifying information.” Users say quality dropped and privacy is gone. Nobody has to be lying. The infrastructure for verifying any of it simply doesn’t exist, so neither side can produce evidence the other could check.

Mechanism: without the key, nothing can be checked

How this family of watermarks works, in one sentence (the longer version is in the earlier post): every time a language model writes a word it picks from a list of plausible candidates, the watermarking algorithm uses a secret key to quietly boost some of them, no single sentence looks unusual, and re-scoring enough text with the same key makes the accumulated bias visible. When I wrote that post, Anthropic hadn’t published its algorithm and I could only infer that Claude’s scheme belonged to this family. The technical note confirms the inference and supplies the specifics: SynthID-Text hashes the most recent tokens (a four-token sliding window in the published configuration) together with the key into a pseudorandom “preference,” then tilts the model toward preferred candidates when several words would do equally well (Nature paper, October 2024). The more of that bias a passage accumulates, the more confidently it can be flagged; the signal strengthens with text length and with how much room for word choice the text offered (arXiv:2609.09604).

Two things follow directly from this design. First, detection needs only the key and the text, with no access to the model, which makes it cheap and scalable. Second, and this is the source of everything below: without the key, nothing can be verified. Whoever holds the key holds the power to verify. Right now, nobody outside the vendors has the production keys, and the vendors also control the detection endpoints.

Watermarked, and still unverifiable

Gap one: the quality dispute can’t be settled. Google reported a controlled experiment in Nature covering roughly 20 million live Gemini interactions: with the watermark on, thumbs-up rates differed by 0.01% and thumbs-down rates by 0.02%. Anthropic reports no impact in internal testing. But outside researchers can’t test the production Claude or Gemini; they can only test open-source implementations. The arXiv paper did exactly that, running the open SynthID-Text release on two open-weight models, Gemma-2-9B and Llama-3.1-8B. On prose (500 open-ended prompts), the watermark’s effect stayed within the noise of simply regenerating with a different random seed, and a judge model scored watermarked against unwatermarked output at roughly even odds. On code (364 programming problems, 10 samples each), Llama’s pass rate fell from 63.8% to 60.7% while Gemma barely moved. The catch: this is an open-source replication. Vendors don’t disclose what strength or parameters production uses, so the watermark actually running in production has never been independently tested. Users who say their code got worse can’t prove it; vendors who say nothing changed have no third-party backing either.

Gap two: “no identifying information” can’t be falsified. Anthropic’s exact words: the watermark “carries no identifying information and can’t be traced to a specific person, organization, or chat.” Mechanically, that can be true: the preference comes from the key, and nothing resembling an ID is written into the text. But the paper puts its finger on the loophole: if the vendor assigns each customer a different key, then trying the keys one by one at detection time and seeing which one matches reveals which account the text came from. Not a character of the text changes; the tracking capability comes entirely from how keys are assigned. One global key, or one key per customer? From the outside there is no way to tell, and the sentence “carries no identifying information” is literally true in both worlds. Users’ fear lands exactly in that blind spot.

Gap three: nobody vouches for detectability. Article 50 requires that outputs be “detectable,” and the paper’s measurements pour cold water on that. At a 1% false-positive rate (one human-written passage in a hundred wrongly flagged), only 39% of watermarked prose was identified at 200 tokens, rising to 56–59% at 400. On code, detection was close to guessing: AUROC of 0.55–0.57, on a scale where 0.5 is a coin flip and 1 is perfect separation. The reason is entropy, meaning how many reasonable choices exist for the next word. Code has little of it: syntax and correctness lock most choices in place, leaving the watermark nowhere to hide. In the paper’s runs, a quarter to a third of code samples came out identical, character for character, with the watermark on or off. Anthropic concedes the same limit: “watermarking is sparser on factual passages where there are fewer choices that can be made without decreasing the accuracy of the text.” So in the settings people worry about most, homework code and factual claims, the watermark is at its weakest.

Set these three gaps against the statute and the mismatch is plain. Article 50 demands marking that is “effective, interoperable, robust and reliable” (the same sentence adds the qualifier “as far as this is technically feasible”), and names no one to verify any of those four words: no accredited auditor with access to keys and detection configurations, no shared evaluation protocol, no requirement that vendors’ detection interfaces talk to each other. The AI Act has general market-surveillance and complaint machinery, but nothing in Article 50 sends an auditor to check a watermark, so whether a given watermark meets those four words rests on the vendor’s own declaration. Anthropic’s detection API is currently open only to “eligible organizations as required under EU law”; the paper records that the interfaces available today compute a score internally and return only a verdict (Google’s open-source detector at least gives a three-way call of watermarked, unwatermarked, or uncertain), never the score itself. In the earlier post I argued that if detection capability goes only to platforms and institutions, the information gap between the detectors and the detected deserves more scrutiny than the watermark itself. A month later, that gap is the default product design. On the vendor’s own declaration every clause can be counted as met, and the legislative goal, a public able to identify AI content, is still out of reach.

Who this is already touching

Content platforms can hardly use the watermark for screening at scale: detection access is gated by each vendor (Anthropic’s API is a private preview that eligible organizations must apply for, and no shared cross-vendor interface exists), the verdict comes with no score to recheck, and a wrongly flagged user can contest the finding only through the vendor’s own tooling; the paper calls for an independent appeal channel precisely because none exists today. Dispute scenarios are worse. A student accused of turning in an AI-written thesis, a freelancer suspected of filing machine copy: the only verdict on offer comes from the vendor’s detector, which is a black box, so in effect the vendor is witness and judge at once. Misinformation policy faces an asymmetry more basic still: anyone intent on deception can use an unwatermarked open-weight model, or paraphrase the output wholesale. The research the paper surveys shows that full paraphrase substantially weakens or removes statistical watermarks, and Anthropic itself says “a complete rewrite where every word is replaced” will strip it. What the watermark actually covers is the everyday output of rule-following users of rule-following vendors.

The paper’s prescription has five parts: publish paired watermark-on/off samples, disclose deployment configuration, create accredited audits with key access, standardize evaluation protocols, and make detection interfaces interoperable. They all aim at the same thing: turning “the vendor says so” into “anyone can check.”

My own judgments, two of them. First, until verification machinery exists, no AI-text detection result, positive or negative, should be the sole basis for a high-stakes decision: academic discipline, an employment dispute, evidence in court. Detectors without watermarks already supplied the precedent: OpenAI retired its own AI-text classifier in 2023 over low accuracy. Watermark detection stands on firmer principles, but the configuration actually running in production has never been independently tested. For now, its reliable territory is good-faith provenance work: vendor self-checks, aggregate research statistics.

Second, the question worth watching in the next regulatory round is not the watermarking algorithm but key governance: how many keys exist, at what granularity they are assigned, who may submit text for a query, and who keeps the query logs. Those answers decide whether the watermark is a content-provenance tool or a user-fingerprinting system that no statute ever named. Article 50 only asks for a “machine-readable” mark; the invisible watermark is the implementation the vendors chose. Anthropic and Google have now both chosen it: Claude’s is on by default, and Gemini’s has been running in the consumer product since late 2024. The real test began the moment the mark went in.

References