Blog

Watermarks, Watermarks Everywhere, and not a Stop to Think? EU AI Act Art 50 comes to town.

Posted on    by
Blog

Watermarks, Watermarks Everywhere, and not a Stop to Think? EU AI Act Art 50 comes to town.

By 27 August 2026No Comments

Blog post by Lilian Edwards, Honorary Professor, CREATe, Glasgow University and Director, Pangloss Consulting

I got an interesting query from Business Insider recently. Anthropic recently announced, as many of you will know, that it will embed watermarks directly into text generated by Claude.  This is intended to fulfil its obligations under art 50 of the EU AI Act (AIA) requiring providers “generating synthetic audio, image, video or text content [to] ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated”. (And incidentally, possibly also its duties under the very similar California Act that also came into force on 2 August 2026).

Inevitably, a developer, Guillaume Meyer, has now released a tool that he says can remove the watermark from Claude-generated text. The slightly less inevitable part is Meyer claims he has done so not to enable fraud, hide copyright infringement or otherwise break the law, but because he doesn’t like the way Claude does watermarking. Anthropic candidly admits that a detected mark only “indicates that the content may have been processed by Claude.” It “is not fully conclusive.” It “does not, on its own, confirm the full provenance of the content”. In other words, it is not a label indicating the text was solely the work of AI. Despite this upfront admission, there has been a serious pushback from small business users and online communities, notably on Reddit, who argue that this approach may over-reach what the AI Act requires. Their concern is that the watermark may label users as pushing wholly AI generated content, even if they  only use Claude as part of human creation of text, and/or editorially review the process. Given the backlash against AI slop, some, such as fiction writers, have been loud about fears that misleading credentials could wreck their reputation.

But what third parties assume from the existence of a watermark (see below) is more a problem for societal norms than the fault of Anthropic, who have been admirably upfront about what they are doing. Watermarking text is much harder than dealing with images or video, where “cutting out” an embedded watermark will tend to leave obvious traces, if not visible then detectable. Text by contrast is continually edited, melded, amended, changed in format, has metadata attached and removed, mainly for quite legitimate reasons. The C2PA standard, which Meyer prefers (as do I) works well for images. It relies on embedding a cryptographically signed file of metadata. But if text is copied from an application or website and pasted into a plain text editor, email, or a different document, all file-level metadata is completely lost.

Anthropic thus  claim that although they are using C2PA for images, for text, it is not really a working solution. Instead they have turned to a solution similar to Google’s open-sourced framework Synth ID, where text is watermarked not by embedding any kind of image or metadata, but by subtly reordering the previously random choices of the Claude model regarding tiny aspects of writing style at token level. These changes are applied systematically, and can be recognised by a detector program, which then “labels” the text as in some part “AI-generated” by Claude.  Given this very different approach to C2PA, it seems reasonable at scale not to make a distinction between distinction between text that is fully AI generated and one to which Claude has contributed. Article 50(2) does say that a watermark is not required where the AI merely provides an “assistive function for standard editing” but this is not the same as saying it cannot be added in such cases. If AI watermarks were to be legally banned except in 100% complete cases of AI generation, then likely the utility of EU AIA art 50(2) –  which is in essence to alert users to the possibility that text they encounter may be “artificially generated or manipulated” – not to say who wrote it – would dwindle to nothing.

I have my own doubts about the worthiness of these kind of labelling obligations anyway –ordinary people will go along with them when they are automatically added to content in the Anthropic/Google style, the unscrupulous will evade them, the free speech digerati will find services which don’t add these labels or remove them. It is always the way, look at the hopeless history of mandatory spam labelling – but that’s for another blogpost. Right now Anthropic seem to be making a good faith effort to comply with actual law in a way that is reasonably effective and commercially practical. (Google did something very similar nearly three years ago without similar blowback, which seems strange). I don’t think the fact that Anthropic are explicitly stating that their solution labels text processed through Claude as  AI-manipulated as opposed to 100% created, gives carte blanche to developers to “ethically hack” their solution. Caveat emptor: if being identified by a Claude watermark would destroy your reputation, better to avoid using it even partially.

So, does Meyer’s tool  destroy EU AIA, art 50 before it is barely out the gates? No. First point – this is hardly new. People have been producing tools to strip out metadata and watermarks forever – often to evade copyright claims but also as in this case for what are claimed to be public interest concerns. I document an example of C2PA, once the gold standard of metadata, already being hacked in high-end phones in my recent EU AIA chapter, at fn 52 (forthcoming, Edwards, Goanta and Urquhart Law, Policy and the Internet, Hart , 2nd edn, this September.)

But more interesting questions arise.

  1. Is the developer (Meyer) in this case breaching anything in the EU AIA?
  2. Are Anthropic breaching art 50 by using a tool which can already self-evidently be subverted?

On the first point, nope.  This isn’t the emphasis of the AIA which is about requiring providers of AI systems and GPAIs to attach these watermarks in certain circumstances. Those obligations apply only to providers, not to third parties who don’t make AI systems or large models.

But that defers the question to other EU legislation, national laws or, most likely, private terms of service. What does Anthropic demand of users in its subscriber contract and acceptable use policies?  Note that the developers of this tool might hypothetically not even have a contractual relationship with Anthropic, but probably they do. Looking at the consumer terms, there’s nothing that on the nose directly applies, but plenty that could be interpreted to include it as a violation, eg “You also must not abuse, harm, interfere with, or disrupt our Services, including, for example, introducing viruses or malware, spamming or DDoSing Services, or bypassing any of our systems or protective measures.” [italics added]. That leads to possible sanctions including having the account suspended, and possibly being sued for breach of contract. Normally I’d say, so what – people take a new subscription and hide their real world identity from lawsuits all the time – but in this unusual case it’s a plausible threat. But I very much doubt Anthropic will want to create (even more) online hate by suing this public interest warrior.

But the second question is the crucial one. Not only can watermarks be deliberately hacked,   as Meyer has shown, but even in non-adversarial circumstances, it’s clear that Synth ID style watermarking will not only over-include but also under-include. Anthropic’s watermark, just like Google’s, will struggle to stick if the text is heavily rewritten, translated, or if the output requires exact unalterable strings as with computer code. So if all textual watermarks can be hacked, diluted, removed, at least sometimes, can any provider really comply with art 50?

Art 50(2) requires providers to “ensure their technical solutions are effective, interoperable, robust and reliable as far as this is technically feasible, taking into account the specificities and limitations of various types of content, the costs of implementation and the generally acknowledged state of the art” [italics added]. So, yes – the law is not completely an ass. The Commission’s Code of Practice for art 50,  which is, NB, advisory not binding, but would very much be taken into account in any enforcement proceedings, has a whole section on what “robust”,  especially, means – see Measure 3.3 which follows standard “take state of the art precautions” lines.

Inter alia it states:

Signatories will choose the considered malicious behaviour as plausible real-world threats based on the type of content, the type of mark and the context of deployment and dissemination of the content. Signatories are encouraged to apply standard cybersecurity practices, such as rate limits, to prevent and counteract malicious use and attacks against their marking and detection solutions. Signatories are encouraged to frequently update their threat assessment to keep up to date with changes in the threat landscape.”

So, I don’t think this kind of watermark remover realistically threatens Anthropic as long as, they have chosen an over-protective, risk-averse solution. Unfortunately, this is exactly what has enraged the Reddit commentariat.

The truth is most watermarks, especially for text rather than images, are inevitably vulnerable. It is yet another arms race, just like all AI detection tools. But most non-nerdy or non-criminal people don’t engage with dodgy off-site apps and the effectiveness of the watermark in general consumer and commercial practice will probably survive. The real question going forward is what use will be made of watermarks, potentially compromised or even absent, by third parties – in disputes about copyright violation, consumer deception and judicial evidence, to name a few. A watermark does not prove who wrote the text, or directed the prompts, or what their intentions are. It is simply provenance data which will be used in various ways. I imagine consumer protection authorities or copyright lawyers will find evidence of someone using tools to remove watermarks quite a smoking gun; lawyers  will inevitably use the existence of tools like Meyer’s  as a useful way to cast doubt on watermarks as any kind of proof. Courts will be told that the absence of a watermark may or may not prove evidence is fake and AI generated, depending on what works for the client. Lots of £££ for forensics experts beckons (as usual) and lots of confusion for low level judges, without a budget for or training in forensics. This is just another tiny milestone in what I have begun to call “the destructive pollution of the infosphere”.

One derivative which I have particularly worried about for some time is how this might affect automated content moderation. In a world where watermarks or metadata are universally added to digital text, will the presence, or even the absence of such, start to be seen as presumptive evidence the content is in some way unlawful/deceptive/fraudulent? Or even actually untrue? That can go bad places: proof of provenance should stay very closely demarcated in impact,  and not extended to pretend to be automated proof of legitimacy. According to a reputable UK parliamentary briefing earlier in 2026, no major platform has yet started filtering by watermarks, as opposed to labelling; but you do wonder when, not if, it will come.

California’s AI Transparency Act, noted above, will in 2027 require large platforms not just to watermark AI generated text but to provide a free tool to detect that watermark. This will obviously be helpful in some circumstances (eg student plagiarism) but will also beckon to more questionable watermark censorship. (Interestingly, the California act will also require platforms with over 1 mn users to refrain from stripping out watermarks in hosted content “to the extent technically feasible” – but it still does not include any criminal provisions for individual developers such as Meyer.

Which leads us to a final interesting point which is whether creating tools to assist removal of watermarks is any kind of illegal hacking. Probably not (every text editor strips out metadata – I use it every time I prepare a Word doc for posting online), but one interesting provision is the UK Computer Misuse Act’s section 3A which makes it a crime to make or sell a tool ’likely’ to be used to commit, or assist in the commission of a hacking (unauthorised access) or Distributed Denial of Service (DDOS) offence. This offence needs intent so will be quite difficult to establish but it is an interesting route to look at, and many countries will have similar laws as it derives from the UN Cybercrime Convention, art 11. The UK’s own prosecution (CPS) guidance asks prosecutors to think about whether there is a public interest in prosecuting under s 3A, something which was aimed at protecting worried cybersecurity researchers. In Meyer’s case, he is at least claiming a legitimate public interest concern; whether the CPS would always believe such claims is another story and would probably depend on whether there was an income stream involved.

The overall point is that we must be careful what uses we make and expect of watermarks.  They provide (some) inevitably flawed provenance data; they do not guarantee wholly human creation, or truth. Just as the registered keeper of a vehicle may not have been the person driving a car at the time of an accident, the law finds it helpful to creates evidential crutches in an uncertain world, but if pushed past their utility they can have antisocial effects.  Think carefully before you throw the first digital stone at an AI-watermarked novel.