An AI watermark proves a model touched your text, not that it wrote it
Both major labs switched on text watermarking within 48 hours at the turn of August. My first thought was that this is very hard to do in a way that actually holds — text has no pixel layer to hide anything in. So I read Anthropic's own documentation on how it works.
It does work. It also breaks in a lot of ordinary situations, and they say so themselves: short text leaves too little signal, heavy editing or translation thins it out, screenshots and format conversion strip it entirely.
But the part I did not expect is the one nobody seems to discuss. A detected mark says the model processed the text, not that the model wrote it. Proofreading, translating and summarising leave exactly the same mark as generating from scratch.
I dictate most of what I write and run it through a model to fix grammar. The thinking is mine. The mark would be there anyway.
Read it the other way and it is no stronger: no mark found proves nothing at all.
So what it catches reliably is one narrow case — pasting raw output and changing nothing. Anyone who edits carefully or runs the text through a second model comes out clean.
I still think it beats having no signal, and Anthropic published the limitations instead of burying them. It just cannot carry the weight people will want to put on it.
Would you want your own writing marked because you asked a model to fix your commas?
Replies
I think this is useful as one piece of evidence, as long as people don't treat a detected mark as proof that the model wrote everything.
WebCurate.co
Exactly. There’s a huge difference between AI touched this text and AI authored this text. The former feels much easier to detect, but also much less useful.