Anthropic Embeds Watermarks in Claude to Detect AI-Generated Text
Anthropic has deployed a watermarking mechanism across its Claude language models to address transparency concerns surrounding AI-generated text. Starting with models released on or after August 2, the system embeds an imperceptible digital signature into generated content that persists during copy-paste operations and may survive light editing. This update fulfills commitments made under the European Union AI Act and applies globally, including usage via third-party cloud providers. Anthropic plans to extend the capability to older models and will release detection tools for external developers. The technology aims to assist the publishing and academic industries in verifying content origins. Recent disputes, such as the agent's withdrawal of support for Jerry Falade's novel Call Me, I'll Hide the Body and Hachette's rejection of Mia Ballard's Shy Girl, underscore sector challenges in detecting unauthorized AI assistance. While watermarks offer a new method for auditing authorship, Anthropic acknowledged technical limitations; heavy editing, translation, or mixing with human text can render signatures undetectable. Additionally, a detected watermark does not conclusively prove generation by Claude, as the artifact can remain present through secondary processes like proofreading. Anthropic joins Google DeepMind, which launched text watermarking via SynthID in 2024, marking a growing industry trend toward embedded content verification.
