Then bad news: LLMs already use randomness in a fundamental way. Each time they go to generate a token, they first generate a probability distribution of possible tokens. Then they pick one randomly according to this distribution. The technique described can be thought of as making the random number generator pseudo random. The output it generates is one of the possible outputs it would have generated before, just now it's deterministic and will generate the same thing every time.
Does the author think he is currently getting T=0 output from Claude? Is he under the impression that T=0 produces the "best" writing?
This entire article just seems so detached from the basics of how LLMs work.
Then why are you using an LLM to write? They're not capable of understanding such nuance. They do pick randomly between two synonymous phrases, they do not use some super smart algorithm to pick the one that sounds the best.
This excuse doesn't hold any water at all - Occam's razor says the author is just super annoyed that his AI writing will be identifiable as AI writing.
I think that was intended, yes.
I honestly can't stand the way Claude writes. This watermark change just makes it scarier.
At the point that you're generating entire volumes of text from Claude you're not really trying to be a sophisticated writer. I don't see how it's going to hurt for it to choose random related words.
Who could have seen this coming???
LLMs have never been the place I've thought to expect any commitment to the craft of writing, to be fair.
Either they false positve on pretty much everything ever written, or the chances of catching a true positive is so low as to be useless.
Basically Cinavia for text, and that often falls over and is easy to remove even when there is megabytes of data streaming over a long period of time rather than 2 or 3 bits per wall of text, let alone what most people use claude for, when there is a strict dictionary and other tight output constraints.
The "watermark" can be trivially defeated, but may be enough to satisfy the letter of the law, and like many people here, I would argue that if you are letting Claude write for you, you've already accepted getting the literary equivalent of turd soup, so the harm is — or at least could be — fairly minuscule.
[1]: https://digital-strategy.ec.europa.eu/en/policies/code-pract...
(FWIW I have a more favorable view than most people seem to of the EU's efforts to at least try tackle problems like this — but predictably, the bureaucratic "solutions" they come up with don't work, but do make things objectively worse)
Total variation distance has been measured to decrease as you scale a model, and that is the primary mechanism "watermarking" as discussed in the Anthropic announcement relies on. It becomes more difficult to reliably detect text as a fixed sample count without tweaking the distribution further. Either way, it's a minor problem that will be addressed over time, compared to the issue of who can detect this without guessing or developing their own sets: providers not releasing a way to detect any such watermarks without going through them makes this entire approach hostile to the public. The EU regulation on this subject is interesting, although again most certainly not the primary driver for these practices:
"1.1.2: Signatories will ensure that AI-generated or manipulated content is marked with an imperceptible watermark, with the exception of very short text. For free-form text longer than 200 tokens, watermarking still needs to be applied, even though it may have lower reliability compared to that of watermarking very long text"
A proper, effective and useful law would have required providers to regularly release datasets to run your own verification on any text released within a fixed interval of time, presumably once out of rotation. Instead, it only talks about exposing an user interface going through their own services:
"Signatories will ensure access to their detection solution through a user interface appropriate for the audience of end-users that may eventually be exposed to the content generated or manipulated by their AI system. [...] Any restriction to the access will be limited in time until more reliable and robust detection mechanisms have emerged and have been adopted as the state of the art for detection mechanisms for the watermarking of free-form text evolves."
Most interestingly, in line with the EU's mass-surveillance program, an alternative solution to watermarking where it may not be sufficient is also suggested, although only optional for now:
"Where appropriate and taking into account potential trade-offs related to privacy and security, as well as scalability challenges and costs, Signatories may implement as an optional supplementary measure fingerprinting or logging solutions for AI-generated or manipulated content which allow for checking whether content has been generated or manipulated by their AI system. For example, direct logging may be appropriate for text content, whereas fingerprinting approaches may be preferable for audio and visual content."
And what I mean by that is that companies that are at the top tend to make anti customer decisions because they have lost the concept that pleasing customers matters as priority one.
“Absurdly and insultingly”? Come on…
I struggle to see how this could possibly be useful unless there's some sort of psy-op going on to trick people into uncritically accepting anything lacking a watermark as not being AI-generated.
Yeah, John. We're all OK with that.
It already fails. It randomly picks between close candidates. To help fool people into believing in intelligence claim, I guess.
FITFY.
I have no sympathy for writers whining about what the AI is doing to 'their' writing. It's only your writing when you write it. There's any easy way to avoid this: don't fucking use it. Use you own brain.
... and the opposite is also true, sometimes it will increase the probability of choosing the "best" word choice. So watermarking makes the LLM quality better then? /s
So, you don't own the generated text, and can't use it freely then. What if I copy paste a section, or rewrite a section of text to my liking? What if I rewrite some lines of code that contains the mark?
Security theater, and vague enough to be used as a weapon against who the government wishes.
I hope it's left off for non-EU customers.
Well, akshwally...
> Interoperability. Providers must implement an interoperability solution for watermark detection such as a standardized API access method, a publicly readable signpost mechanism embedded in content, or participation in a consortium detection solution by February 2, 2027
Reminds me of printer tracking dots.
The author has expressed a preference. Assume that there is a sequence of tokens, such that it is considered the absolute best by the author. This particular method of watermarking makes it less likely to generate that sequence, by definition.
I feel their argument would have been clearer and stronger if they had spent more time exploring the alternatives, and whether these alternatives would be just as effective. It is trivially easy to remove invisible tokens.
Like it or not, there is a public good to being able to identify AI generated content, and a small degredation in quality is tolerable in my opinion.
I don't think anybody has to worry about this issue though. Manual writing, coding, and proof reading continues to be an option. Where AI output is nothing to be ashamed of, the tools are available. For everyone else, there will be LLM providers that ignore EU law.
Try running an llm like qwen 3.8 27B in Q8 locally with an intentionally very low temperature setting, it will write like a caveman crossed with a robot. You may find that an extremely literal output does not look pleasant to read for humans.
Watermarking is bad not just because of the principled stance that your tool should not be working against your own interests (the passionate argument in TFA), but specifically because it lends credence to the idea that AI detection is a valid and possible thing to do perfectly.
As technologists of course we know "oh well yes but with some confidence interval we can detect AI token bias across a large corpus of text." To JimBob in charge of publishing your paper or reviewing your PhD submission, all he knows is "anthropic says AI detection is possible so this 30% chance your paper was written by AI means you've plagiarized." Do you really think you're winning the argument with the certified, law-approved plagiarism detection machine? No, you're not, and your career is over.
It's irresponsible to develop watermarking because it is not anywhere close to a perfect science, but it will be treated like one by people with the power to ruin your lives. Even if you've never touched AI in your life, your paper is going through the "maybe it says you cheated" box, and you better hope those dice don't come up snake eyes.
> The difference between watermarked and un-watermarked text will not be distinguishable to readers
https://www.anthropic.com/news/claude-text-watermark
Which is to say, it does not actually meet the EU AI act requirements which require transparency to humans. Not to mention that if the detection requires access to the base models, it makes anthropic the only entity who gets the say on if a piece of text comes out of Claude. Anthropic is both the player and the referee here.
If there is one takeaway you should have from this fiasco it is that you should be wary of using tools that doesn't serve your needs and your needs only.
Nanny state nonsense indeed.
As a non American/non European resident all I see from the Europeans are 0 contributions to software progress at any large scale while they surely do a lot of crying and huffing and puffing and demanding. Lots of complaining and rule making but not a lot of creating is a bad look.
Even so, I don't think it will stop here. Once this is in place, the next step is to put more and more identification into the AI generated content; might as well pack it in, it's not that bad, and if it is they won't admit it. There's no way for anyone to check. And your argument will still be technically correct but missing the point.
Cognitive surrender.
(1) The behavior that is approximately what you describe is not "fundamental" (though it may not be something you can disable on some hosted providers), it is an option that is not fundamental (and with runtimes where you have full control can be either disabled or tuned in a large number of manners), and
(2) The actual behavior that is approximately what you describe already usually involves use of PRNG (with a user or harness supplied seed), not a true RNG; the change to do watermarking isn't going from RNG to PRNG, it involves adding an additional set of constraints on token generation on top of the existing ones, which inherently compromises quality.
On the other hand, LLMs are forced into picking some likely-ish word, and then have to build the rest of their response to retcon that choice into making sense.
Even good human writers would probably struggle with this constraint. It would be like someone interrupting your writing to tell you the next word MUST be such-and-such, and then you have to try and make it work as best you can first try, without going back to edit. The result would probably be a little clunky. (Maybe it’s impressive LLMs write as well as they do.)
I'm skeptical that anybody generating LLM text is really all that concerned about optimal word choice. Or even particularly good prose. But let's pretend that person exists.
If that person tried, say, an open model and that same model with watermarking applied, I'd be eager to hear their thoughts on the prose quality. Especially if they built an experiment harness and rated a few hundred blinded examples and found a measurable difference.
But getting this upset in advance of any demonstrated problem? It really seems to me like the point isn't the point
But it feels to me like you would need a hell of a lot of text to bury even a simple account ID. The nudges they are talking about are of the order of a handful of bits over several hundred words, I think?
You can't assume that because if that was the case he'd already know what sentence to write, because that's what that means.
The notion of a best sentence requires a final cause, an end to write to. By their very nature that's not how LLMs work, so you can't 'degrade' them on that front. They can't lose a property they didn't have.
Variable name perversion incoming...
How would that work? Claude appending " written by AI" to each of its messages? That would both be impractical and useless.
They will need to dodge around the EU requirements but it will probably just come down to an alternative method to watermark or a contractual assurance you won't mis-represent the source of the text.
No and no. I am not sure I agree with his point but I know he is not ill-informed on either of these points, because I mentioned them to him a couple of days ago.
It's still possible to use Claude to proofread - highlight grammatical, flow, structure, logic errors and make simple suggestions for you to pick and choose or adapt as you wish. No watermarking will flag your text. No flaw accusations of LLM authorship will haunt you. All will be fine.
But if you want an LLM to rewrite your text, that's (a) not proofreading, and (b) should be flagged as LLM generated ... because it is.
And if you copy-paste the answers from LLM, I think it's only fair the end result gets flagged. You're not writing it yourself.
A point of confusion for me, however: is every watermark unique? Is every algorithm for watermarking going to vary amongst models and amongst model versions? Will each model publisher keep this watermarking as a trade secret, that they alone can detect? If so, this can't scale! How do you detect "JoeBob 4.3 LLM" output? By querying every single model's watermark-detector? And if they all work by re-running the model and using tokens anew? That is extraordinarily wasteful.
If a watermark is not self-evident, or universally detectable, then it is no good. Take, for example, US currency. The security measures are published and well known. Any count-out room in retail has a big poster indicating how you can detect authentic US bills. Nobody has to accept non-US currency in the US, and so the only authenticity you need to worry about is your US bills alone. LLM watermarking has none of this in common. Currently sounding like a shitshow, if you ask me.
How does that follow? AI-generated text is already not a perfect emulation of human writing. There's lots of room to affect it laterally without changing the level of quality.
As I understand it, LLMs with temperature >0 can select from many possible outputs. All they're doing is limiting the possible outputs to ones that contain this pattern. I don't see any reason why the quality of that subset should be lower than average. The very best outputs will likely be eliminated, but so will the very worst.
In fact we know it's not that good because we can often tell Claude's writing apart from human writing.
Could it have been equal or better with slight variations in wording?
The slipper slop argument is too lazy to address directly. Argue A is bad because A, not because A might become B and you’ve got good arguments against B.
https://chainofbranches.com/conversations/2/branches/20/
I’m not convinced it’s possible. A good nights sleep and a notepad in a quiet room still feels like the state of the art toolchain for writers.
On deeper tech stuff, like this utterly nonsensical misunderstanding of watermarks… yeah, classic case of a guy who is smart, and has lost the ability to realize when they’re not knowledgeable in a domain.
The "problem" is that seeing the watermark doesn't mean that the person claiming to be the author didn't make extensive changes to the output of the LLM, or that the LLM wasn't simply the final editor of something that the author had put a lot of work into.
> Cognitive surrender.
I don't know what this means. It's just drama. Don't let the LLM write for you and this is not a worry. I'm not worried about the poetry of LLM output being subtly adulterated.
I don’t see how this follows? Tokens are chosen randomly. If you choose tokens with a different RNG in the same distribution, you’re still getting equally good or bad tokens.
Well, I suppose it's nearly the opposite of that experience, upon further review. But for some reason, that's where my head jumped.
If you get caught uploading watermarked media without the clear label, you're in big trouble, mister.
> One of my fundamental problem with this is that no two synonyms carry the exact same meaning. “He leaped at the chance” and “He jumped at the opportunity” are very similar sentences expressing the same general sentiment, but they are not the same. The exact words we choose when writing matter.
Doesn’t make sense at all in light of the actual approach, they’re just choosing a different RNG. It’s not like they’re corrupting it by flipping words.
Should add I don’t support the watermarking and requiring it is idiotic.
I wonder if it would even get flagged in that case, because wouldn't the probability distribution of a token when the LLM is suggesting an edit to your writing be different than the distribution of that token once it is in the context of the text it's editing?
Presumably you could use the same reasoning trace, run multiple generations, and get different outputs (if the temperature is >0).
But now I’m interested in playing more with Cowork or Claude Code/Codex for prose writing to see if the set of tools there affects outputs at all. I guess you might need a more custom “writing” harness.
You be a human who's brain shifted into LLM mode (chainneling Markov?).
Or perhaps you're an LLM impersonating humanity.
I often wonder how much LLMs are just mirroring our own brain's patterns.
You're mixing up two claims here, and only one of these is kind of true. Yes LLMs do internally plan ahead in a way that is emergent rather than strictly part of their architecture, so that part of your claim is true. The way you word it by saying they are "coalescing the probabilities of a range of tokens at a time" is poetic sounding jibberish though. What's actually happening is one distribution output for the next token computed from a hidden state that implicitly encodes where the text headed.
Your claim that if an LLM does happen to pick a token "th" instead of "tw", then the LLM isn't stuck with that decision is entirely false for autoregressive LLMs which is what all of the frontier models are. Whatever an LLM picks as its output token is final, it has no ability to undo that token selection and it must continue on the basis of that choice. It can't go back on that decision and revise the output.
If you're interested in this, Anthropic has a summary of a very technical paper on this topic that mostly deals with this issue with respect to poetry:
https://www.anthropic.com/research/natural-language-autoenco...
Mathematically, a long chain of conditional probabilities is equivalent to a single probability over the whole range. But computationally, for that to work out, the computation for the first probability needs to somehow consider all the downstream probabilities depending on it, which obviously isn't how autoregressive language models work. They can pack in as much downstream computation as their neural architecture allows for, which is quite a lot.
Suppose in some context you have three equally plausible conpletions after "Be": "tween a rock and a hard place", "twixed he stood there" and "lieve he can fly". To model this probability distribution of the whole sentence, the next token "tw" needs to appear at 2/3 probability and "lie" at 1/3. After "tw" would be a 1/2 chance of "ix" and a 1/2 chance of "een"; after "lie" would be a 100% chance of "ve " and in any case the rest of the sentence after that would be 100%.
The model needs to somehow "think ahead" to know those are the possible completions. For example if "lieve he can swim like a dolphin" was another equally plausible completion, that first token would need to be 50/50 instead of 67/33. So the computation of the first token somehow needs to encode the fact that the guy thinks he can fly but not swim, even though it doesn't become relevant in the output until several tokens later.
In practice this probably happens to some degree but definitely doesn't happen perfectly. To perfectly model the first token's probability distribution, it would have to include knowledge of the entire distribution of all possible outputs, which is just not happening. So it approximates. Surprisingly, the approximation is good enough to produce language.
You can see this breaking down in the seahorse emoji incident from last year. When you ask the model if there's a seahorse emoji, it first completes "Yes," as if a few tokens later it's about to produce a seahorse emoji. But when it actually gets to the token that would produce a seahorse emoji, it can't because there isn't one. But it's already outputted "Yes, the seahorse emoji is" and can't just go back and change that to "No, there's no seahorse emoji." Some models would try a few times and then say there isn't one or a system error seems to be making them unable to produce one, other models (including then-current ChatGPT) would loop forever with ensuing hilarity.
Here’s an example: I had asked Claude for some music recommendations in a certain style. Part of its output was:
—
*Long journey tracks*
Clinic — “The Return of Evil Bill”
Guided by Voices — not really, wrong band
Silver Apples — “Oscillations”. Proto-everything, deeply repetitive, hypnotic.
—
So at some point there, the next token produced was “Guided” or “Guide” or whatever, and then because it can’t go back, it had to correct itself after the fact.
Reasoning/CoT have helped a lot, but I feel like small versions of this still happen all the time.
Human writing is like 90% editing.
The point I made (quite briefly) is that watermarking is only feasible because for good writing it is necessary to use T>0, or the writing will never explore a more creative choice, and that at T=0 you don’t even need a watermark to spot LLM-generated text.
The point he is making is consistent with this, isn’t it? Either you allow temperature to drive creativity, consistently in a way that can be influenced and analysed, or you adulterate that process for the purposes of meeting a corporate/legal directive, in a way that is proprietary and obscure. These are ethically distinct approaches, and since he disagrees with the EU objective he comes down on one side I guess.
Me, I don’t care about the hypothetical enough.
Not least because I think Claude writes depressingly badly and I doubt any steganographic change will enrage me less.
Wouldn’t having that be enough to eventually reverse engineer the key?
Seems like he really likes to use LLMs and is worried that quality will be degraded. But he will never demonstrate such degradation scientifically, we don’t have anecdotes even.
Working backwards: if it is possible to confirm 100% confidence that a chunk of text is LLM output, then it is "PD until proven otherwise". How can a human reliably assert human authorship of their source text? When all watermark tests fail? Is that proof of humanity now?
If a human proves human authorship, and LLM watermarking tests positive, then is that going to be considered a "derivative work" or not? What if there is an applicable license for the source work, such as "CC-BY-ND" that prohibits derivative works?
This has not been court-tested, and I expect that it will need testing at that level before we can have any assurances.
But what about the general idea that they can watermark results to tell where they came from. The next step is tracking down which user got a result. I hate both of these things. Must everything we do be tracked? Next altering wikipedia results so they can tell who looked at the page or something?
I'd like "the best answer" from an llm and don't want to be tracked, but this isn't for me, it is for them. I understand llm results are already using a varying statistical input so they aren't always the same. But I really hate watermarking and likely tracking too.
Perhaps it would be useful to publish examples of samples with/without watermark. I'd suspect that the variability from simply sampling repeated times would dwarf any semantic differences you'd detect with the watermark.
Related: if you don't have a limit on sampling (top-K or top-P), eventually you'll hit one of the really unlikely tokens by chance and then the model will switch to Japanese because the most likely completion after a random Japanese character in the middle of an English sentence is more Japanese writing, not a reversal back to English.
Reasoning tokens with tool calling tell the model to loop on a one phase of a question and call a tool to indications completion when done.
Related, but not the same thing.
So we train a second copy of Claude to work backwards—reconstruct the original activation from the text explanation. We consider an explanation to be good if it leads to an accurate reconstruction. We then train Claude to produce better explanations according to this definition using standard AI training techniques.
Incentives to train a pathological liar. There's no baseline so can only catch out the worst of the lies/errors. Anything (including fabrications) that passes our filters is reinforced?Removal may come down to changing every third token to a different one.
He repeatedly states that choosing "the best word" is the most important thing to him. I don't know how you reconcile that with creativity itself, let alone probabilistic sampling.
As for it being "uncopyrightable" if it were the output of an LLM, I think computer people are making a very aspie interpretation of a single decision. I think it's more that the LLM (and thereby its owners) cannot itself hold a copyright on its output, that output has to be touched by a person before it is copyrightable. A particular view from the top of a mountain can't be copyrighted, for example, but a photograph of that view can be.
I'm not sure it at all precludes a "robosigning"* sort of situation, where machines generate output, hired temps sign and claim that output, and immediately sign it over to the people who hired them (as a work-for-hire.) Copyright is stupid, artificial law, not logical.
-----
* https://www.mortgageauditsonline.com/what-are-robo-signers/
You are projecting that onto me, and I cannot tell you how comically poorly aimed it is.
For the watermark to be detectable, the text needs to be like 75% AI generated.
If you have an LLM “touch” one section of the article, it’s not gonna be detectable.
But also, no one really knows as they're closed.
When I wrote this week about Anthropic’s announcement that all Claude models, worldwide, would soon begin “watermarking” everything they generate, including text, to comply with this EU regulation, we were left to speculate how this was going to work, because Anthropic offered not even a vague description of how it would work — despite the fact that the title of the announcement was, absurdly and insultingly, “How Claude Marks AI-Generated Content”.
My initial speculation was that maybe they’d hide invisible non-printing Unicode characters in the text. Just spitballing. Turns out that’s not what they’re going to do. What they’re going to do is apply a form of steganography, where the choice of words (or other token output) at inference time will leave fingerprints that can later, maybe, be detected probabilistically.
I initially guessed “invisible characters” not because I didn’t think of the semantic word-choice technique, but because I was a fool who took Anthropic at its word in their description of what they would do. Their original support document claims:
When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
They say “imperceptible” and “doesn’t change the meaning, quality, or readability”. Their words. Not almost imperceptible. Not slightly changes the meaning, quality, or readability. That made sense to me, because that’s absolutely what I want — nay, demand — from any tools I use personally. It’s unacceptable for a tool to sacrifice an iota of clarity, coherence, meaning, quality, etc. for the purpose of embedding hidden clues within the text to suggest its provenance. That’s what I would and will demand. And Anthropic’s (original) support document unambiguously claims that’s what their system will enable. So if that were true, I couldn’t see what was left other than hiding invisible characters within the text.
My error was believing Anthropic that their system wouldn’t adulterate and corrupt the semantics of the text their models generate. That is in fact exactly what they plan to do. I should have my head examined for believing a single word of a document titled “How Claude Marks AI-Generated Content” that doesn’t explain, at all, how Claude marks (or will mark) AI-generated content.
Yesterday, on an entirely different website than the original “How Claude marks AI-generated content” article (the one that didn’t explain anything at all about how it works), Anthropic published “How Claude’s Text Watermark Works”, which does actually explain in layman-accessible terms how it’s going to work. I will return to Anthropic’s new highly euphemistic and slightly misleading description below.
There’s a bunch of research on this topic, some of which I have also linked to below. But the very best description of the general idea behind the technique is an interactive essay by James Padolsey, “How AI Text Watermarking Works”. It’s a wonderfully cogent read, and the interactive elements splendidly illustrate the main concepts. A+ work. If you have any interest in this at all, I dare say you must read — and play with — Padolsey’s piece.
But here’s my stab at a layman’s high-level summary. If you toss a coin N times and note the results, you can determine with a degree of certainty whether the coin is fair or biased. LLMs are, in their popular incarnations, non-deterministic. Ask the same question of the same model and you often get at least slightly different answers. Maybe the same meaning, but different phrasing. At each decision point for generating the next token, the model makes a choice. With these semantic watermarking techniques, they make different choices for some tokens based on word lists that could be called “green” and “red”. At each decision point, they’re a little more likely to pick a word from the green list than the red list. That doesn’t mean they never choose words from the red list. Just that they’re less likely to than they would if the adulterated marking technique weren’t in place. (Same way that a crooked 51-49 coin will still land “wrong” side up 49 times out of 100 on average.)
Words or word phrases are sorted into the green and red lists deterministically on the fly, at each “next token” generation point. So sometimes a specific word will be on the green list, and other times it will be on the red list. Someone with the secret key can determine which list a word will be on at each token generation point (which is how the watermarking is detected); those without the secret key cannot. This means there will never be a list of words that Claude prefers or eschews.
With coin flipping, the higher N is — the more times you flip — the more confident you can be that the coin is fair or biased. So too with this semantic watermarking. The more words in the text, the more accurate the analysis will be that the text was generated by a specific AI model or not. With too few coin flips, you can’t achieve any confidence at all regarding a coin’s fairness. With too few words (or tokens), there’s no way to achieve any confidence whether a string of text was AI-generated or not.
Given a string of text to examine for signs of a specific watermarking system, if there are more words tagged as green and fewer tagged as red than would otherwise be expected, the text can be flagged — with some degree of confidence — as having been generated, or merely modified, by the AI system that applies the specific secret-key watermarking system. The amount of confidence in the determination will obviously vary, significantly, based on the size of the text string and randomized weights given to words on the green and red lists. But only Anthropic will be able to determine if text was seemingly generated by Claude, and Anthropic will only be able to detect the watermarks that are applied by Claude. Claude can’t detect the hidden watermark signals generated by, say, Gemini, and Gemini can’t detect the hidden watermark signals created by Claude, because each implementation is predicated on secret keys held only by the LLM provider.
One of my fundamental problems with this is that no two synonyms carry the exact same meaning. “He leaped at the chance” and “He jumped at the opportunity” are very similar sentences expressing the same general sentiment, but they are not the same. The exact words we choose when writing matter. I want any LLM I use to choose the very best, most precise words at every single decision point. An obvious constraint that I accept is time and computation. Within the constraint of executing inference quickly, and at a certain cost per token, I want the best words. This constraint matches human writing. I could surely write a better column by taking longer to write it. I write with a sense of how much care I should put into every word and punctuation choice I make. I take more time with certain paragraphs, sentences, or even individual word choices when my gut feeling says I should.
In other words, these are necessary trade-offs. These factors are all in my interest: speed, cost, quality. Ideally I would like perfect writing, at instantaneous generation speed, at zero cost. None of those things are possible. Computation is not free of charge (and cloud-based LLM inference with leading models is actually expensive). Inference is not instantaneous. And great writing, whether natural or artificial, can only approach perfection.
The idea that anything other than my needs should factor into the generation of text for me is patently offensive.
This isn’t just about text one might generate with the intention of passing it off as their own natural work. This isn’t even about LLM proofreading of work written by hand. Anthropic is saying that all new Claude models are going to adulterate every single bit of text longer than 200 tokens (~150 words) they generate, including everything it presents to its users to read. So even in a private conversation between a user and Claude, which will never be read by anyone other than the user, Claude will begin making word choices in the name of marking its output in statistically predictable ways rather than maximizing clarity and precision.
Even today’s so-called frontier models are already decidedly lacking in lucidity. Claude, ChatGPT, Grok, et al. are “better writers” than most humans and produce better prose than the median human. But: no shit. Most people are terrible writers. The “average person” is pretty stupid and half of all people are stupider than that. And there are many smart, interesting people who are miserable writers. So as impressive as LLMs are, the bar is low. The best writing I see come out of these models is worse than anything I would choose to read for pleasure. And now Anthropic is saying they’re going to make it worse, on purpose, for purposes that do not benefit me in any way? Even if only slightly worse?
Get fucked.
Speaking of objections, the relevant EU regulation motivating all of this, “Code of Practice on Transparency of AI-Generated Content”, is red-tape nanny-state pipe-dream nonsense. Here’s Ben Thompson’s summary from a paywalled Stratechery update this week:
- The regulation applies to text longer than 200 tokens.
- The provider must mandate in their terms-of-service that users not remove the watermarking.
- The solution should be robust in terms of evading “typical processing solutions” like screen shots, scanning and OCR, copy-and-pasting, translations, etc.
Taken literally, compliant LLM terms of service must forbid users from rephrasing the output from models that comply with this regulation, because the word choices are the marks. But it’s not the European Union that is trying to impose their absurd, impractical, witch-hunt-fueling regulation on the entire world. That falls on Anthropic.
Complying with this, particularly with regard to text, is only going to create problems for honest users. Dishonest users attempting to pass off AI-generated text as their own writing (students, employees, whoever) will simply circumvent detection through non-compliant AI paraphrasing tools.
James Padolsey — whose interactive visual explanation of how these schemes work I linked to above — explains this in a post titled “Anthropic’s Weak Watermarks Appease a Weak Law” (which, if it rings a bell, I linked to in a standalone post earlier today):
The same thought that led to this law could have applied to calculators at the time of their inception, had their outputs revealed themselves through artefacts. Thankfully, a sum borne of the brain is treated no differently from one produced by a calculator. Likewise with spellcheckers. To make assistance suspect only once the tool becomes capable enough to compose a whole sentence is not a principled boundary. It is a moral premium placed on difficulty itself.
Anthropic has nevertheless chosen a blanket, model-level implementation that appears broader than the law’s minimum requirement. That may be convenient compliance engineering, but it discards distinctions the law expressly attempted to preserve. The result is a signal broad enough to implicate harmless and assistive use, yet fragile enough to be removed by a motivated person through substantial recomposition. It risks concentrating suspicion on ordinary and assistive users while remaining weakest against deliberate deception.
Padolsey is the creator of Declaude, a delightfully simple web app that allows you to “Paste in AI-flavored text and get the same content back as plain prose”. Declaude’s original purpose is cleaning the saccharine Claude personality stink from text (whether it was created by Claude or any other LLM), but, if Anthropic persists in its stated plan to begin adulterating all text Claude generates, Declaude will also serve as a copy-paste single-extra-step way to eliminates those marks. Declaude is interesting and useful already, but it exemplifies how ill-considered and futile this EU regulation is when it comes to prose.
Google has a watermarking system in place that they call SynthID, which they apply to AI-generated images, video, audio, and text. I’m concerned in this article only with text. With multimedia, embedded watermarks can be metadata within files, and truly not affect the experiential quality of the work when viewed or listened to. With text, we are talking about the actual words that are chosen. From the “AI-generated text” section of Google DeepMind’s own description of SynthID:
We’ve expanded SynthID to watermarking and identifying text generated by the Gemini app and web experience. Large language models generate text one word (token) at a time. Each word is assigned a probability score, based on how likely it is to be generated next. So for a sentence like “My favorite tropical fruits are mango and…”, the word “bananas” would have a higher probability score than the word “airplanes”. SynthID adjusts these probability scores to generate a watermark. It’s not noticeable to the human eye, and doesn’t affect the quality of the output.
In a group chat, a friend of mine quoted the above, and I responded that if a chatbot wrote “My favorite tropical fruits are mango and airplanes”, I’m pretty sure I’d fucking notice. Another friend then responded with this:
Days later, that still cracks me up.
But Google’s absurd description puts the lie to their own claim that it isn’t noticeable, and it serves to show just how little regard the people behind these generated-text fingerprinting schemes have for the actual craft of writing. Of course bananas has a higher probability score than airplanes, because airplanes aren’t fruit. But what about pineapple? Should the sentence complete to “mango and bananas” or “mango and pineapple”? That’s a good question, and the only acceptable answer for why an LLM should choose bananas instead of pineapple (or coconut, or guava, or papaya...) is that it has determined that it’s the best fit for the intended meaning, tone, and sentiment of the text. Not because bananas is on the watermarking “green” list and pineapple is on the “red” list, even though pineapple might be the better fit. Google’s own supposedly jocular description of how SynthID works in fact captures how the scheme perverts the text it generates.
They’re saying you won’t notice because if it only chooses bananas over pineapple for these fingerprinting purposes, well, they’re both tropical fruits and who cares. But it’s utter nonsense that the difference is “not noticeable to the human eye”. The semantic difference between banana and pineapple is just as noticeable to the human eye as the taste of the two are to the human tongue.
If it did produce “My favorite tropical fruits are mango and airplanes”, it’d be incredibly stupid, but it wouldn’t be offensive because we’d all recognize that something completely off-key happened. What’s offensive is that with a system like SynthId in place, where the fingerprinting decisions are motivated by a secret key, we have no idea whether it completed to “mango and bananas” because bananas was determined to be the best next token, or because bananas is in the “green” bucket of words. It calls every single word choice into question.
Here’s a paper published in Nature where Google’s team behind SynthID published their work, after putting it into production with Gemini (née Bard):
We analysed approximately 20 million watermarked and unwatermarked responses and computed the thumbs-up and thumbs-down rates (both as a fraction of the total number of thumbs-up and thumbs-down feedback received). We found that the thumbs-up rate for the two models differed by 0.01% (with the watermarked model being higher); and the thumbs-down rate differed by 0.02% (with the watermarked model being lower). We found both of these differences to be statistically insignificant, and well within the 95% confidence intervals.
From this experiment, we conclude that over a wide variety of real chatbot interactions, the difference in response quality and utility, as judged by humans, is negligible. Subsequently, non-distortionary SynthID-Text has been productionized and is currently watermarking responses in Gemini and Gemini Advanced. To the best of our knowledge, this evaluation represents the first systematic watermarking investigation of its kind within a large-scale production system.
To this I say:
Gemini/Bard’s thumbs-up/thumbs-down buttons are not a good experiment for evaluating the effect on quality. If a chatbot tells me “My favorite tropical fruits are mango and bananas” instead of “mango and pineapple”, I’m not going to give the response a thumbs down because of the fruit it chose. I’d give it a thumbs down if it said “airplanes”, yes, but that’s a strawman. (The paper in Nature even uses “My favourite tropical fruit is ...” as an illustration, but in the paper, the only four next tokens considered are, in order of probability distribution, mango, lychee, papaya, and durian. No airplanes. And, conveniently, in the paper’s example, the “winner” of the watermarking “tournament” just happens to be mango, the one that would have been selected as the best if the watermarking weren’t in place.)
A “difference in response quality and utility, as judged by humans” that is “negligible” does not mean imperceptible. What they really mean is that it’s only slightly worse and that everyone is either too stupid to notice or too indifferent to care.
It’s widely considered that Gemini is behind ChatGPT and Claude in quality. Perhaps the fact that they’ve put SynthID-text into production is one of many reasons why. I personally agree that Gemini’s prose is inferior. Maybe the use of SynthID has nothing to do with the fact that I, along with the general public consensus, consider Gemini to be a second-rate chatbot — but in that case, maybe it’s the fact that Gemini is a second-rate chatbot that makes the difference “negligible” when Google started mixing in SynthID-motivated tokens in its results. It’s a lot more likely that your restaurant customers won’t notice that you replaced your regular coffee with Folgers Crystals if your regular coffee is second-rate to start with.
Now, finally, back to Anthropic’s new “How Claude’s Text Watermark Works”, published yesterday. I have some comments.
To summarize:
We use a method of watermarking that does not have any practical impact on the quality or content of Claude’s outputs;
The difference between watermarked and un-watermarked text will not be distinguishable to readers;
Translation: Specific words do not matter and we don’t think anyone reads anything closely.
- Nothing is added to the text and there are no hidden characters;
This would have been worth clarifying at the outset.
- Watermarking won’t be specific to Claude. As of August 2, the EU requires AI providers serving its market to mark AI-generated content. Other major model developers have signed the same Code of Practice and will be implementing their own watermarks.
No other AI provider has stated that they will apply such marking, adulterating all generated text, outside the EU.
Take the sentence “The weather today was cold and…”. The next word is very unlikely to be “sugary.” But it is quite likely to be “overcast” or “grey.” Under most circumstances, it doesn’t matter much to the reader which of these latter two words the model ultimately chooses — the meaning of the sentence is largely the same either way. In cases like this, the choice is settled by a random number.
Arguing that grey vs. overcast “doesn’t matter much to the reader” is the crux of my argument that this entire endeavor is a perverse adulteration of what it means to write — or to read. That it’s subtle in some ways makes it more perverse, because it’s sneaky.
In internal testing, we’ve seen no impact of watermarking on the content, level of creativity, or readability of Claude’s text. In the SynthID-Text paper, which introduced the technique we use, Google DeepMind tested this impact by serving a model that used watermarking to a portion of their Gemini traffic and comparing thumbs-up and thumbs-down ratings. They found no statistically significant differences from the unwatermarked model. And in a controlled study, human raters comparing watermarked and unwatermarked answers side-by-side saw no difference in quality.
See above for my argument that this thumbs-up/thumbs-down data is absolutely worthless in evaluating whether the SynthID-style word-bias watermarking makes text worse. By definition it must make text worse, unless the underlying LLM model’s scoring is wrong, because the nature of the watermarking algorithm requires it to sometimes increase the probability of selecting a worse word choice and decrease the probability of selecting the model’s best choice. It’s only a question of how much worse. What Google’s thumb-counting data shows is only that it isn’t so much worse as to make Gemini users click the thumbs-down button.
Watermarking doesn’t change the meaning or experience for the person reading it, but if you wanted to check after the fact whether the text was likely generated by Claude, the watermark allows you to do so.
No, it does not. Because the entire scheme is tied to secret keys held only by the AI provider, it only allows Anthropic, not “you”, to check anything.
When Claude proofreads text written by a person, what it gives back has generally only been lightly edited; because nearly all the words are the person’s, there’s very little (if anything) for the watermark to attach to. Depending on the length of the text and how heavily Claude has edited it, those changes might not be enough to make Claude’s involvement detectable. The more Claude writes, the more decisions it has to make, and the more space there is for a watermark.
Translation: No one can ever again use Claude for proofreading their own prose unless they’re willing to risk that the whole thing might be flagged as having been generated by Claude.
For example, once the model has written “2 + 2 =”, there is a very clear best choice for the next token (if the model is completing the sum, there isn’t an answer that’s equally as good as “4”; if it’s talking about George Orwell’s Nineteen Eighty-Four, there isn’t an answer that’s equally as good as “5”). The “nudge” of the watermark wouldn’t be applied here. For the same reason, code — which in very many cases has to be exact — has generally less watermarking than some other forms of text.
Having said that, in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code. But by definition, it will have a negligible effect on the actual code produced.
Translation: We value precision in programming code; we do not in prose.
And it is exceedingly rich to cite George Orwell’s Nineteen Eighty-Four, approvingly, in the context of justifying a text adulteration scheme premised on the notion that specific words do not matter. I mean what the actual fuck? Orwell!
Lastly, as to why they’re doing this:
We’re implementing watermarking to comply with the EU AI Act. Anthropic, along with several other major AI model providers and around 190 total signatories, signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026. This requires AI system providers to use methods of “marking” AI-generated text. We’re applying watermarking globally at launch because we don’t yet have a durable way to scope it by region.
This, from a company that the Financial Times just reported is weeks away from an IPO with an intended valuation of $2 trillion, which would make it one of the 10 highest-valued companies in the world — as of today, placing it at #7, between TSMC ($2.2T) and Broadcom ($1.9T).
This leaves us to believe that one of the following must be true:
It’s perfectly reasonable that a technology company valued on par with Amazon and TSMC is technically incapable of complying with an EU regional law only within the EU itself.1 Not a cause for concern at all.
Anthropic is in over their heads, wields shockingly little control over their own tech stack, and their imminent IPO is likely to be remembered only as a new high-water mark in the manic global AI bubble.
Also, what happens if another major global market makes it unlawful for AI to secretly watermark generated text?
From an OpenAI support document titled “Provenance Signals (Content Credentials, SynthID) in OpenAI-Generated Content”:
Consistent with our commitments under the European Commission’s Code of Practice on Transparency of AI-generated content, our goal is to expand provenance signals to all modalities including text, so customers and developers have clear ways to meet their own transparency obligations as standards and tooling continue to mature.
There’s a lot of wiggle room in this brief statement, and it could just as well mean that OpenAI models will only adulterate text with fingerprint markers when users or developers ask for it. Or that it will only be mandatory for users in the EU. If I were at OpenAI I’d go hard on this and publicly say that ChatGPT will never watermark text it generates unless you ask it to, and that if you want tools that secretly work behind your back without telling you how they work to flag your words in ways you can’t see, go ahead and use Claude.
Three papers on ArXiv:
I will admit that while I’m profoundly offended by the idea of personally using tools that attempt to leave such watermarks in text they produce or touch, the mathematics behind it are fascinating.
Michael Lopp, at Rands in Repose, “RIP Claude”:
As a human who has had to wrangle with EU regulations in the past, I am abundantly clear what’s involved in the laborious bureaucratic process. I can guess what threats Anthropic is facing. However, this is a tone-deaf, clumsy, and alarming opening salvo in their watermark strategy. [...]
My writing is my work, and Anthropic’s current strategy is aggressively writer-hostile.
Jeff Gamet, “Anthropic’s Claude Watermark Is Akin to an AI Poison Pill”:
To be clear, the watermarking is embedded in pretty much any text Claude touches. Along with text Claude generates, it also applies to text it processes, such as proofreading and summarizing. I expect we’ll see too many inaccurate accusations of using Claude to write documents where the content was human-written, but AI-proofread.
The watermarking sticks with documents through copy-and-paste, too. Imagine copying text from a blog post or email only to have what you wrote tagged as potentially AI-generated. In fact, that could very well happen with this post. I personally write all of my content without AI tools, but I copied the quote at the top of this piece directly from Anthropic’s website. Does that mean what I wrote here will show as AI-generated? If they used their own models to generate or edit what I quoted, then the answer is very likely “yes.”
One of the papers published at ArXiv I cited above claims that such watermarking even persists when an article of text originally generated in English is translated into German.
Secrets are the poison here. When only Anthropic holds the secret keys that both produce the watermarking and perform the probabilistic detection of those marks, we’re all left to wonder. To wonder if what we’re reading is secretly watermarked, what we’re quoting is secretly watermarked, and whether what we ourselves are writing will be unjustly accused of being AI-generated based on secrets we don’t know and can’t see. Poisonous is exactly the right word.
Or should I say toxic? Or airplanes?