To comply with the EU AI Act Anthropic will watermark its AI output. This means that, technically, texts generated by Anthropic’s models should now be more detectable as being AI generated. Although this sounds like a good idea, it will only make the internet more synthetic, will create an oligarchy in AI detection, and would still fail to provide certainty in AI detection.
Watermarking in technology, both visible and hidden has been used for decades. These could be a symbol on a picture or information in the metadata of a file. The problem with AI-output, however, is that it is plain text. You cannot watermark the text as it is easily removed by the user. So, Anthropic will now use a method called SynthID, whcih is already used by Google Gemini since 2024, to watermark its output. The idea is that the AI model will prefer certain words over others when generating text in the generative loop. This creates a lexical pattern that detection tools would be able to trace.
AI-output already suffers from a phenomenon called homogenisation. As AI models are predication algorithms, they tend to produce texts that are more centred around a grey average rather than being rich and diverse. SynthID will only increase the repetitiveness of this output, using model specific words to create the watermark.
It is estimated that half the internet is AI generated. The more people will be exposed to synthetic texts, text that has been generated by AI, the more they will use similar language in their writing. If you read a lot about ‘delving into the matter’ and ‘fostering ideas,’ chances are pretty high you will use those words in your own writing. If the words ‘delve’ or ‘foster’ are part of the watermark system of a specific AI model, your writing might trigger a false positive with an AI detection tool. As a student it will become more difficult to prove your innocence if the institution claims its detection tool makes use of the AI watermark.
A solution to this problem could be to rotate the preferred words of a model to create a watermark. This would result into slightly more diverse texts and reduce the copycat effect by AI users. However, this would also create a dependence on AI providers for detection tool companies. These companies need to be notified regularly by AI providers about the changes, otherwise their products become useless. And still, no certainty can be given whether a text has been, partly, written by an AI model.
Another problem is that nefarious students will find models that will not watermark their text, use (modified) open models, or use tools to scramble their AI generated texts. The ‘watermark’ only provides a false sense of protection. Even Anthropic warns that the SynthID system is not water-tight: “it doesn’t confirm whether the text was human-written.” They also admit that editing the AI generated text might circumvent their ‘watermark.’ It seems Anthropic just complied with EU regulation without providing something that is actually useable.
Educational institutions do well to (still) abandon AI detection tools despite these watermarks and to find other ways to assess students’ understanding and skills. Not only will detection tools make these institutions more dependent on the AI companies, they will also provide a false feeling of certainty whether something was written by AI. The introduction of AI has forced educational field to reconsider evaluations. Sticking to methods from the past is only water under the bridge.




