What is AI watermarking?
The EU AI Act’s article 50 specifies that LLM providers need to make any AI generated content identifiable as such. This can be done with a technique known as watermarking. While it applies to audio, image and video content as well, we’re going to look at watermarking for text here. The main AI companies have decided to apply this technology globally, as doing so will be much easier than trying to identify which outputs are being used in the EU. The upshot is that it will be possible to take a piece of text and feed it into an algorithm controlled by the LLM provider, and find out with a reasonable degree of certainty whether the text was written by a particular model.
You might come across some tools that claim to be able to find watermarks in the whitespace and remove them. That won’t work for this type of watermark. The watermark is contained in the actual choice of words, not in how they are presented. That means you won’t be able to remove the watermark by copy-pasting the text somewhere else, doing anything with the formatting or even copying the words out by hand. As long as the words stay the same, the watermark remains intact. Even changing the text to some degree or mixing it with other text not written by that model won’t completely stop it from being detectable - the identification algorithm will still be able to identify that some part of the text originated from the model with some level of certainty. Of course, the more you change the text, the less confident the algorithm will be, until at some point it won’t be able to detect it at all. But the idea is that if you’re changing the text so substantially that it’s indetectable it will cost you enough work and effort that it might not be worth using an LLM in the first place.
How does it work?
As we mentioned, the watermark is contained in what words the model selects. It is artificially biased in a way to pick words in a certain way, which is statistically detectable, but doesn’t make the output of the model any worse.
Fundamentally what an LLM is a system that, given any set of words, predicts what word is most likely to come next. Let’s say you have the following sentence
AI constitutes a revolutionary new
Here are some plausible ways you could finish that sentence:
technology 0.41 | development 0.26 | paradigm 0.22 | discovery 0.15 … banana 0.02
The numbers are the scores that the model might assign to each word - the higher the more likely. Without watermarking the LLM would do one of two things - it might just pick the highest scored word and return technology. Or it might add a little bit of randomness and occasionally choose a slightly lower-scored option, say paradigm shift. This randomness is called “temperature” and can serve to make the output less robotic and predictable. With watermarking we want to change the way the word is selected without making the selection significantly worse than it otherwise would be. If it consistently ended up picking low scoring words like banana in this case that wouldn’t be any good.
Instead what happens is that there is a tournament of sorts between all the possible choices, where higher scoring words are entered more often, giving them a greater chance of winning. But each round the winners are determined by a secret method that only the LLM provider knows. If anyone else knew it, it would be easy to obscure the watermark. In this example, let’s imagine the rule is that we go for the word with the highest number of the letters L M and N. This would lead to development winning more often, even though technology is a better word. But because technology was entered into the tournament more often it still has a better chance of making it through to the end than a bad word. The detector would then check how often words with more Ls Ms and Ns appear even when they are not the best choice. Of course the actual rule would be nothing so simple. Instead it generally uses some kind of hash function.
How does this affect me?
Watermarking in itself changes nothing about how you are or aren’t allowed to use LLMs - it simply makes it so that LLM written text becomes identifiable. That means the main consequence is that if you are using LLMs to write text and you don’t want your audience to know that you are, they will now be able to tell with a much greater degree of certainty. It is generally advisable to be transparent about where and how you are using AI regardless. Trying to hide the fact you are has always been a bad idea, and now it has become an even worse one.
Whether using AI generated text is acceptable naturally depends on the context. Nowadays many of us use it in most things we write - often we’re even expected to. But in some areas the mere suspicion of AI use can provoke a vicious backlash, as a youtube content creator recently found out the hard way. It will be interesting to see how cultural norms on this develop over time, and where it will and won’t be considered acceptable to use AI.
That is a debate that will continue to evolve over the coming years. Watermarking is simply a mechanism that will allow us to establish the facts around it with more certainty. And this goes both ways. While it can out people who did use LLMs it can also establish that someone didn’t use a particular LLM, at least not without substantial subsequent modifications. The AI-detection tools that have been used up to this point have not been 100% reliable and false positives can have severe consequences for people’s academic careers. Watermarking techniques are probably not going to make this completely bulletproof either, but the fact that a secret key gets embedded in the text at the point of creation should be able to give us substantially more confidence in these verdicts.

This blog post was written by Hilary Roberts, Veritas_fox CTO. Book a free discovery call with Hilary to find out how can Veritas_fox help you with your AI Governance and Compliance.



Comments