Notes ·
An AI Watermark Can't Tell You Whether Anyone Was Thinking
Joshua MorrisSteve Hargadon wrote an essay called "Go Ahead and Watermark" that gets at something I have been thinking about as AI becomes a normal part of how people write. Anthropic recently announced that future Claude models will generate text containing an invisible statistical watermark. It is not a hidden character or identifying tag. Instead, the watermark changes the source of randomness Claude uses when choosing between otherwise reasonable words, leaving a statistical pattern that can be detected across a long enough passage.
Anthropic says it is implementing the watermark to comply with the EU AI Act and will initially apply it globally because it does not yet have a durable way to limit the system by region. Steve's reaction is basically: fine, mark it. His argument is that the watermark answers a much less interesting question than people seem to think it does. It can tell us that a particular tool was probably involved in producing the words, but it cannot tell us whether the person publishing those words had anything worth saying.
Steve uses photography as an analogy, and I think it is a useful one. A modern camera handles exposure, autofocus, white balance, stabilization and an enormous amount of image processing. We don't therefore conclude that the photographer had nothing to do with the photograph. The photographer still decided where to stand, what to point the camera at, when to press the shutter, what to keep and, more importantly, why the photograph was worth taking in the first place.
Writing with AI can work the same way. I use AI constantly, but the useful part for me is not asking it to manufacture an opinion and then putting my name on whatever comes back. I bring the question, the experience, the disagreement, the technical understanding, some strange connection I noticed, or simply the thing that has been bothering me enough that I want to understand it better. AI can help me investigate it, challenge my assumptions, organize what I am thinking, find weak spots and sometimes help me express something I already understand but have been struggling to put into words.
That is considerably different from typing "write me an article about AI" and publishing the response, but a watermark cannot tell those situations apart. Anthropic is unusually clear about that limitation. Its own documentation says the detector can determine the likelihood that Claude was involved with a piece of text, but it cannot distinguish between Claude generating the work and Claude heavily editing something a person wrote. The watermark is also weaker when Claude has fewer meaningful word choices available, which is why proofreading, factual passages and code may contain much less of a detectable signal.
So if a detector finds the watermark, what have we actually learned? We have learned that Claude was probably involved with the text at some point. That is useful provenance information, but it tells us almost nothing about the intellectual contribution of the person whose name is attached to it. We still have to read the thing.
We used writing as a proxy for thinking
The part of Steve's essay I found most interesting is his argument about education. For a long time, producing a polished essay was expensive enough in time and skill that we could use the finished essay as a rough proxy for several other things. If someone could produce good prose about a subject, we assumed they probably understood the subject, had organized their thoughts and had done much of the intellectual work required to arrive there.
AI breaks that relationship because someone can now produce polished prose without understanding very much at all. That is a real problem, particularly in education, but I think we make a mistake when we only look at that side of it. The inverse is also true: someone who understands something deeply but has difficulty turning those thoughts into polished prose can now remove a large part of that mechanical barrier. An AI watermark does not tell us which person we are looking at.
Steve writes about his father, who had an extraordinary career in college admissions and lived surrounded by books and ideas but never wrote a book himself. He wonders whether AI might have helped someone like his father get ideas into writing that otherwise remained mostly spoken. I like that way of looking at the problem because we tend to discuss AI writing as though everyone begins with the same ability to convert thought into prose and AI simply gives some people an unfair shortcut. People do not begin in the same place.
Some people think much better than they write. Some people speak much better than they write. Some are working in a second or third language, and some have disabilities or processing differences that make the mechanics of writing more difficult. There are also people who can produce beautiful, technically perfect prose without having anything particularly interesting to say. Writing ability and thinking ability overlap, but they have never been the same thing. AI is making that distinction much harder for us to ignore.
Provenance still matters
None of this means I think provenance is useless. If someone generates a fake statement from a politician, publishes synthetic news, impersonates another person, or creates media intended to deceive the public, knowing something about where that material came from can be extremely valuable. The transparency requirements in the EU AI Act are concerned with exactly these kinds of problems, including deepfakes and AI-generated material presented to the public as information about matters of public interest.
I would actually like better provenance systems for digital media. C2PA and cryptographically signed content credentials are interesting because they can tell us something about where an image or document came from and what happened to it along the way. Anthropic is using C2PA credentials for supported files created or processed by Claude, which makes considerably more sense to me than trying to infer provenance after the fact from how something looks.
Where I get uncomfortable is when provenance quietly turns into a judgment about authorship, honesty, competence or intellectual effort. Those are different questions. "Claude was probably involved in this text" does not mean "this person didn't write this," and it certainly does not mean "this person didn't think this." The absence of a watermark cannot prove the opposite either. Anthropic acknowledges that sufficiently rewriting the text can remove the signal, while another model may use a different watermarking system or none at all.
If schools, employers or other institutions start treating the presence of a watermark as evidence of wrongdoing instead of evidence that a tool was involved, I think we will end up in a very predictable arms race. Detectors improve, tools for removing the signal improve, people find new ways around both, and we spend an enormous amount of effort trying to reconstruct a distinction the watermark was never capable of making in the first place.
The expensive part of writing has changed
Steve makes another observation that I think gets much closer to what is actually happening: fluent text used to be expensive. Producing pages of competent prose required enough time and skill that the ability to produce it had economic and social value of its own. AI has made that particular part of writing dramatically cheaper, and institutions that used fluent writing as evidence of competence are now discovering that the proxy no longer works the way it used to.
What AI has not made cheap is curiosity, experience, judgment or being right. It has not made knowing which question to ask cheap. It has not made understanding a complicated system cheap, and it certainly hasn't made having something interesting to say cheap. What became cheap is converting some combination of those things, or sometimes none of them, into fluent language.
That distinction is going to be disruptive because a lot of systems relied on the cost of producing polished prose as evidence that the expensive thinking must have happened too. I don't think the answer is to pretend that relationship still exists and build increasingly elaborate detectors trying to restore it. We need better ways to evaluate the thing we actually care about. Does the argument hold together? Are the claims supported? Does the person understand what they published? Can they defend it when challenged? Is there actually anything interesting being said?
Those questions take more work than checking for a watermark, but they are much closer to the point.
Mark it, then read it
I don't have much of a problem with Anthropic putting a watermark in Claude's text. Go ahead and mark it. Tell me Claude helped edit something I wrote, helped organize an argument or generated some of the language. That is information, and in the right circumstances it may be useful information. It just isn't a verdict about the person who published it.
The danger is treating evidence that a tool was used as evidence that the person using it contributed nothing. We don't normally make that assumption about cameras, calculators, spell-checkers, IDEs, search engines or most of the other tools we have accumulated around intellectual work. AI is more complicated because it participates directly in composition, and I don't think we completely understand yet what we gain or lose when we outsource portions of that process.
There are certainly people outsourcing the thinking along with the writing. There are also people using these tools to investigate ideas more deeply, overcome barriers, challenge their own thinking and communicate things they might not otherwise have been able to communicate very well. A watermark cannot tell us which one happened.
We are still going to have to do the annoying old-fashioned thing and judge the work.