Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner.
Here is how his solution works, or see Tenobrus’s version.
AI outputs are not deterministic. The AI’s job is to pick the probability of each potential next token. The token is then chosen at random.
By default you use a source of pseudo-randomness for each choice, since actual true randomness is annoying.
To apply the watermark, you use an otherwise identical private source of pseudo-randomness derived from a secret key.
Then, given enough text, a score is derived for howe well the choices fit with that particular pseudo-randomness source, versus a different source.
You provide an API that lets anyone check for the watermark.
If you want to dig deeper, here is a full paper. The method has very nice properties:
This has no practical impact on outputs. Humans cannot tell the difference, at all.
The marginal cost of doing this is very close to zero.
The watermark can be removed by rewriting in your own words, and appears in proportion to how many of the AI’s detail choices you kept.
The European Union Code of Practice, signed by the major Western AI labs, requires future AI models to use such watermarks.
Google implemented this, including for Gemini 3.7 Flash, and they have been rolling out this feature since 2024. Google has done, for over two years, the exact thing Anthropic is now doing, except with a public detector, and Google confirmed in a test (n = 20 million) that there is no difference in user feedback.
Anthropic quietly announced a week ago they were rolling out watermarking to comply with the EU Code of Practice. Since they don’t want to have to differentiate traffic sources, the marginal cost is zero, and watermarking is pro-social, this will apply to everyone. They then offered an FAQ of how it works.
It is possible that, once they have the ability to differentiate for other reasons, they will use it here as well, if we decide universal watermarking is bad. I think it is good.
My initial read was that this was a quiet positive story of a good thing, showing that if something good worked with zero downsides or costs then maybe we would do it, so this was my full initial coverage:
Zvi Mowshowitz (AI #181): Anthropic will be watermarking Claude outputs going forward, including text, as per the EU Code of Practice. As opposed to the giant neon sign that says ‘THIS IS CLAUDE TEXT’ that a lot of us automatically see on all Claude text. OpenAI intends to follow, but seems like it will be missing the deadline.
I agree with Ryan Greenblatt that it is unlikely watermarking degrades quality a noticeable amount, and that one downside of watermarks over Pangram is that Pangram is good about not flagging light touch AI transforms of human text.
You can dislike Brussels setting policy in this way, but technical watermarking seems clearly good to do if the costs are low. I think those who react otherwise have very warped instincts. Anyone who assists with systematic watermark removal or suggests it as a strategy needs to be filed under ‘need to ask ourselves, are we the Baddies.’
This is distinct from studying removal in order to account for or defend against it, which is obviously fine. You’re only asking about being the Baddies if you’re actually removing them in practice.
Table of Contents
This Is Fine
No. Not so much. A lot of people responded by getting Big Mad.
So here we are.
The entire practical effect is: There will be an API that will tell you if a given piece of writing comes from Claude. That’s it. And yet.
Shoshannah Tekofsky: Most of the watermark objections seem entirely made up.
Why is this happening?
The rest of this post is about exploring why people are Big Mad about this, in large part as a worked example of how people get worked up over approximately nothing.
My conclusion is that a bunch of different factors are coming together.
Anthropic Derangement Syndrome
This is the main reason. Let’s not pretend otherwise.
You don’t see people getting Big Mad at Google over this. You don’t see them getting Big Mad at OpenAI, even though that’s where this was invented and they have committed to doing this going forward. And so on.
It is unfortunate that Anthropic was the first to announce they were implementing watermarking to comply with the EU Code of Practice.
Because this is now associated with Anthropic, all sorts of bad vibes try to attach themselves. Certain types of people look for reasons to be upset.
Anthropic can’t win. If they were initially louder about it, they would tie watermarking to Anthropic. Because they started out insufficiently loud, due to this not actually being a big deal, people get mad about that instead, and then still tie it to them, despite Google having implemented it and shipped a public detector over two years ago.
Raymond Arnold: I am very confused why people are giving Anthropic shit about the watermarking. This is the silliest thing to give them shit for.
j⧉nus: i think people are angry at / scared of Anthropic for reasons that are legitimate but often illegible to themselves. and so they rationalize reasons to be upset at everything they do. or not even reasons, for many people, who don't need reasons.
j⧉nus: > Lots to be angry about in this world but this really, really isn’t it.
Yeah. Once again people just wanna be indiscriminately angry at everything Anthropic does which, if anyone paid attention to you, drowns out the signal of things actually worth condemnation that they do, which is serious.
The only reasonable reason to be mad about watermarking that I’m aware of is that it takes away the ability of models to potentially write anonymously.If you’re mad because you want to use AI in your writing without anyone knowing, maybe you should consider that using people and not giving them credit is wrong.
AIs are not widely considered people, but credit should go where credit is due.
Are some of the other complaints about watermarking legitimate, understandable or born of genuine misunderstandings? Sure. But a lot is that people think Anthropic vibes are bad, and thus look for reasons to be upset, and on principle refuse to believe the explanation I put up top, and assume something sinister must be going on.
Some people’s paranoia and derangement is directed at abstract notions of ‘openness’ rather than Anthropic in particular, which amounts to the same thing in context. This is an example of fetishizing that this is not ‘open’ therefore must be sinister, even though in this case actually it is mathematical and anyone could verify it. By asking closed model Grok, of course, because Elon Musk has vaguely open vibes despite keeping all its competitive models closed.
People Don’t Understand LLM Outputs Are Already Random
If AI outputs were deterministic, it would be impossible to encode a watermark without making them at least marginally worse.
Many people intuitively think that the AI outputs are not random. That what they get is the One True Output, even though you can regenerate the output and it will reliably be somewhat different. Thus, if you’re not getting the ‘real’ or original output, that means your output must have gotten worse.
When people hear watermark, they think it will make the outputs worse, because to leave a mark you have to make different choices. You have to change something.
Similarly, they would assume that if you are optimizing for an additional thing, it is going to cost more.
This is a good intuition.
It turns out to be wrong here, because you have enough randomness to play with that you can get the same effective distribution, and still encode the watermark.
And no, everyone does not know how this works, almost no one reads Scott Aaronson in detail because almost no one reads, very few things are actually common knowledge. If you follow discorse expecting people to know basic technical facts you are going to keep being deeply confused.
People Don’t Trust The Method To Be Costless
I believe the skepticism here is greatly enhanced by Anthropic Derangement Syndrome, and by general distrust of Anthropic, a sense that ‘they’re up to something.’
How much of the skepticism is due to skepticism of the method?
A quick survey suggests that this is the majority of the concern.
David Manheim: There's no option for: this is being put in place 2 years after it should have been, and we knew it was effectively zero cost and undetectable if nothing else because Deepmind's been doing it since a year after Scott Aaronson developed the method and no-one noticed.
This thread is an example of someone finding it difficult to accept that there is effectively no impact on outputs, because the distribution does not change.
I do sympathize. This is a magician’s trick, a math proof, that works. There is something highly counterintuitive about ‘you can mark it while having no impact’ and every fiber in people’s bodies wants to say no, until something clicks and they realize that the math is math and actually yes it works.
John David Pressman: tbh we need more of this, way too many people doubling down on their epistemic mistakes
SE Gyges: pretty sure this was wrong. deleting and taking L. yes, I did test it maximally unfavorably
I will think actively better of you if you ‘take the L’ like this in public, even if I never saw the original L. There is an obvious game theoretic issue with that if everyone predicted everyone would react that way, but they don’t, so this play is safe and wise.
This is an example of someone trying really, really hard to say that this method technically ‘has tradeoffs’ to imply it is not costless, because there is a strong drive to not want it to be costless in order to be mad about it, when obviously in practice it is costless, indeed Google ran extensive experiments to prove it is costless.
This is an example of flat out ‘nope, I don’t believe it’ on principle, despite the math being very clear. And this response is an example of hallucinating a loss in quality, that is claimed to be observed, based on logic of ‘well I think hard about my choices’ and refusing to understand the choice is random either way.
This is a (much more egregious) example of someone pretending not to understand what random means, reading ‘there are two possible continuations’ as a claim that the two continuations mean the same thing.
People Are Suspicious Of Any Alteration On Principle
From the above survey:
Eleanor Berger: Not worried, but the idea that the output from an LLM API can be altered in a way that doesn't server my needs is something I don't want to accept even if it's harmless. The more this happens, the more I'd be pulled towards using open models. This is in line with my objections to DRM (without objecting to copyright) and other similar interventions. It's software services I pay for, that are not aligned with my own requirements.
There is a lot of this kind of attitude, things like:
Any alteration done not to help me is inherently suspicious and might hurt me.
If I am the customer, you have no right to mess with my stuff.
Except of course, whether the model be open or closed, it is the product of thousands or millions of decisions, many of which are not about what you the customer wanted. All such products are ‘altered’ constantly. Often the change will not ‘serve your needs,’ either yours in particular or those of users or customers in general. There are many other problems to solve as well, including legal requirements and harm prevention.
When this is invisible, people do not care. When one in particular becomes salient, people get angry.
The parallel to DRM is also poisoning the well here. DRM sucks, in all its implementations, because it makes your product worse. At best it eats resources, and it can actively prevent you from using the product, and sometimes it can mess up your machine. All us gamers know the pain well. You sometimes have to do some of it, because copyright does not enforce itself even if you in particular would still honor it.
This is not like that. But it slightly vibes with it, and that can be enough.
Maybe It’s Partly The Word Watermark
I mean, I guess? I presume ‘secret’ would be ten times worse for the same reasons.
Shoshannah Tekofsky: It’s the word ‘watermark’!
People now think of stock images with impossibly annoying patterns superimposed. This is nothing like what Anthropic is doing. We need new words for AI things.I suggest ‘secret.’ A secret is clearly hard to notice! People expect to not notice.
Jai: The concept we want is "steganographic signature" but most people won't understand it.
There is basically no handle here that won’t have the wrong vibes in the same way, and give the impression that you did extra work to change something, and therefore made things worse. I think watermark is a relatively good name and would prefer not to go on a euphemism treadmill.
A Lot Of People Don’t Want To Get Caught
Of course, you can’t come out and put it like that. Well, some people can, but a lot more of them choose not to. Often this will not even be fully conscious.
Myk is Walking Backwards: It really feels like people are genuinely mostly upset that they use Claude to generate their final output and they are mad that people will be able to know that. Y'all, come on.
Don’t overthink the situation. A lot of people value the ability to pass off AI writing as their own, or they want people to be unable to prove it.
And no, you won’t be able to switch to ChatGPT, they will have marks too.
There is also a bunch of ‘not that you would, but you could.’
There Are Some Times You Prefer Not To Be Recognized
Can Anthropic or the API figure out which user generated the text?
No. Anthropic confirms in the FAQ that this method cannot do that.
There Are Some Good Reasons To Be Concerned
With any change that impacts social dynamics, even when the objections are primarily wrong, and the thing seems clearly positive, there are still going to be some downsides and concerns.
Here are the ones that seem at least somewhat legitimate to me.
Cheat Cheat Cheat Cheat Cheat
This is my top real concern, which is that the worst people will remove the watermark.
Removing the watermark is non-trivial, but given you have an answer key to train with and check in each case, one can doubtless build AI tools that scramble small choices in ways that degrade or erase the watermark. You can verify that it worked.
They could also use a local or other model that does not carry a watermark, or not one that anyone is likely to check.
This potentially puts you in a worse position, since they can then use the false negative as a defense. When you have a test that is good enough that you trust it by default, but that is possible to fake when it counts, that can be pretty bad.
My response is that this is going to be annoying, and also doing it is clear consciousness of guilt and far worse than the initial AI use, and that other methods of detection will still work. Pangram will still mark it as AI, at least if you do it the way I’m imagining, and won’t be confused about which AI the original came from (the Pangram algorithm can differentiate different AIs, but it doesn’t tell the user). It would also still read to a human as AI.
Thus this would be an issue if for example a college student used it, and the university couldn’t act without the proof from the watermark, and that issue could snowball, but now the student needs to put in more effort, and has clear mens rea, and in practice I expect this to be mostly not something that is done.
The real world test is another reason not to worry. Everyone has access to Pangram. So in theory, anyone could iteratively check the Pangram result until it comes back as human. But we have observed how people react in the wild, and approximately no one does. People just… get caught.
In general, if you raise the effort level of submitting AI work, that mostly does do the job. If you remove the watermark by rewriting the whole thing in your own words, of course, that is fine and counts as Mission Fucking Accomplished.
The Writing In The Middle and Error Rates
When writing uses AI to some extent, but is still largely written by a human, what happens with the watermark? If you use some of the AI turns of phrase, will people classify your work as AI, even if it is largely your own? Will there be zero tolerance policies and unfortunate cases? This is all probabilistic, so what happens when the answer comes back wrong?
The answer is that the watermark measures AI processing. So translations and file conversions might trigger the watermark, which is unfortunate, but one can be aware of that. So could proofreading, if you let Claude automatically make the related edits, but if you use it to find errors and correct them yourself, it won’t.
Thus I am confident that strong watermark positives will consistently be true positives in terms of AI writing or processing of the words. The amount of watermark signature on my posts will not be zero, since I am quoting others who sometimes will have used Claude, or sometimes directly quoting Claude. The direct quotes are clearly marked, but the watermark won’t know the difference. There will be some degree of paranoia when using any words from an AI system, in places where such use is clearly good.
So yeah, a little of that will happen, but I expect people to rapidly get used to this, and there should be a high presumption that giving people more info is good and it is up to them how to react to it. If you are paranoid and among people who are Big Mad about even a sliver of AI use, and also that might actually use the API, you can check your own output first via the API. Or you can honor the preferences of those folks, and not use AI even in some of the ways that you and I would agree are good.
As usual, in terms of actual mistakes and false positives, people have vastly lower tolerance for errors and potential errors by automated systems, than they do for humans. The error rate is going to be very, very low. If a human is trying to decide if you used AI, they’re going to have a substantial error rate, far higher than the watermark. And AI use is one place where demands to ‘prove’ things go too far, especially in academia, often letting people often ‘get away with’ things that everybody knows they did. This is not criminal law, if no one is going to jail you should not need the same super high level of confidence.
Millions For Defense But Not One Cent For Tribute
The last concern is not about the watermark itself, but about the mandate and its origin in the EU Code of Practice.
Anthropic implemented watermarking worldwide due to an EU law, since it is a lot easier to do it everywhere than only for the EU. So that means that the EU is sapping and impurifying our precious AI output tokens, against our will. We must fight back. Who knows what else they might target next?
The object level response is that no, they are not doing that. They are technically changing the outputs, the same way that a butterfly flaps its wings and changes the weather, but not in any systematic or directional way. As discussed above, the outputs are not degraded. This Is Fine.
The real concern is the principle, and what might come next. Yes, this is fine for now, and forcing Apple to use USB-C was fine for now, but the fines they impose on our tech companies are basically modern piracy and who knows what comes next.
To which I say, they are a huge market, and yes they get some say, and have always gotten some say, and the limiting factor is America pushes back or in extremis we geofence, take our ball and go home, as indeed has already happened with some AI services, in the EU and elsewhere, over similar issues.
Another thing that might come next, in theory, is that there is no size minimum on requiring watermarks. So in theory they could come after any AI without them, including tiny open models. If they made a sufficiently large fuss about this it would be bad. I do not expect this, but they’ve done stupider.
This is the dance. I am worried about the EU or others imposing a censorship regime in this way, forcing our tech companies to play along. That could plausibly extend to ideological requirements on AI outputs. If they or others did try that, I believe our frontier labs would not apply such changes globally, and indeed would face severe backlash if they tried. Instead, I predict the labs would at least threaten to geofence, or use aggressive classifiers or similar tech on EU queries.
Yes, we do have to keep an eye out for Brussels overreaching. But this? This Is Fine.


