
Anthropic is responding to user concerns about its new AI watermarking plans, which were revealed as a compliance measure for the European Union’s AI Act. The company updated its support page earlier this month to disclose that Claude model outputs would be marked as AI-generated. Following a wave of discourse from users, Anthropic published a detailed blog post on Friday to explain the technology and its limitations.
The company states that watermarking is intended to help meet EU regulations and is being applied globally at launch because it does not yet have a way to restrict the feature by region. The technology will first be applied to new Anthropic models, with older models receiving the watermark in the coming months.
Anthropic claims it hasn’t seen an uptick in cancellations since the watermark was announced, despite reports from some customers who said they canceled Claude subscriptions over the feature. OpenAI did not immediately respond to a request for comment from Business Insider.
Related: Manhattan rents hit record high as vacancies plunge
The AI lab cites a 2024 Google DeepMind paper to explain how it intends to embed proof of AI in text without ruining a chatbot’s human-like tone. The process involves word choice decisions that are often settled randomly. For example, a phrase like “grinning child” might use either “joyful” or “cheerful,” with one option selected by a random number generator.
Anthropic describes its watermark as a different type of random number generator. If a user has a “key,” they can detect whether the text follows Claude’s subtle patterns of choice. The company says it will offer an API tool that will allow users and third parties to check text for Claude’s watermark themselves.
The limits of detection
Watermarks do not include hidden characters, weird fonts, or secret text. They are simply the result of word choice. Anthropic emphasizes that the difference is not distinguishable to readers.
Related: 72-Year-Old Woman Pulls Herself Up at 80
Using its key, one can only answer the question “What is the likelihood this was partly written by Claude?” The key points back to Claude’s involvement rather than AI’s involvement in general. Longer passages are easier to recognize as Claude-generated because the model makes more word choice decisions.
Anthropic notes that “factual passages” and AI-generated code will receive fewer marks. The wrong word choice in code could jeopardize accuracy or cause the software to fail to run correctly. The company also addresses concerns that the watermark implies users aren’t the author of the text or code. It states that the watermark appearing indicates Claude processed the content or file and doesn’t change a user’s rights or ownership under its terms.
This technical approach to distinguishing AI writing relies on statistical probability rather than obvious markers. Because the method is subtle, light editing probably won’t remove the watermark completely. A complete rewrite where every word is replaced will break the watermark, though it is arguable whether the text can still be described as AI-generated in that case.
Leave a Reply