Anthropic details how Claude's EU-mandated text watermarking will work
Anthropic published a blog post on Friday explaining how it will watermark Claude's generated text, a practice it had announced earlier in the week to conform with the EU AI Act's Transparency Code. The post answers questions that have stirred debate since the announcement, including whether the mark can be hidden by editing and how it applies to code.
The company said it will use the SynthID-Text approach that Google DeepMind outlined in 2024. When Claude makes "low-stakes choices" — such as using "overcast" rather than "grey" to describe the weather — the model can build a pattern "undetectable to the reader, but is detectable to anyone who has a key that encodes it". Anthropic also said it plans to release a watermark detection API.
Users pushing back in forums and on social media: a Reddit poster saw the watermark as a conspiracy, while another stated that anyone wanting Claude without it wants to "lie to people". According to Business Insider, "dozens" of users on X claimed they have cancelled their Claude subscriptions over the change.
Anthropic did address the Waterproof. It said the object would not impact quality: "To a reader, a watermarked response is indistinguishable from an unwatermarked one." Light editing "probably won't remove the watermark completely", while a full word-for-word rewrite will. For text only proofread by Claude, "nearly all the words" remain the human author's, leaving little for a watermark to attach to. The approach also differs from style-based AI detection tools, which look for repeated telltale phrasing like "this isn't [X], it's [Y]".
Code is a special case. Because the model must produce something that runs, it cannot always choose among equally valid options. Anthropic says in the places with an arbitrary choice such as comments inside the code, the watermark can apply, "but by definition, it will have a negligible effect on the actual code produced".
What this changes concretely: for Claude users, it will soon be possible to verify whether a text was AI-generated, without any visual cue in Claude's output. The company does not be alone — "other major developers have signed the same Code of Conduct and will be implementing their own watermarks". But practical boundaries are clear: lightly edited proofread text may carry no watermark, and a rewritten text escapes detection. That sets a reasonable, not absolute, accountability.
Comentários
A carregar a conversa…
Inicie sessão para escrever um comentário. Iniciar sessão