· Solveion · Perspectives · 6 min read
The watermark sorts people, not text
Three futures are usually proposed for AI watermarking: normalisation, an arms race, or everyone running local models. All three are already happening, to different people. The evidence says the arms race is over, and what the mark now measures is not honesty but sophistication.

We wrote recently about Claude embedding an invisible watermark in everything it generates. The obvious follow-up question is where this goes, and there are three futures people usually propose.
Normalisation. Everyone’s text becomes machine-touched — Copilot rewrites the email, the assistant tidies the memo — so the mark ends up on nearly everything and stops meaning anything.
An arms race. Labs harden the watermarks, a market of removal tools works to defeat them, detectors work to catch the removers, and the cycle continues indefinitely.
Local models. Anyone who cares runs open weights on their own machine and simply never produces a marked artifact.
The usual framing asks which one wins. We think that is the wrong question, because all three are already happening simultaneously to different populations. The better question is what a watermark is actually worth once all three are true at once.
The arms race is not a race
Start here, because it changes everything downstream.
An empirical evaluation published this year tested whether watermark evidence is forensically usable. Two of the leading text watermarking schemes lost their mark on 100 per cent of previously-detected samples after a paraphrase pass. Google’s SynthID managed 98.3 per cent removal. A separate benchmark found a single pass through ChatGPT dropped every method tested below 30 per cent detection.
The baseline numbers before any attack are arguably worse. The same evaluation reported false-negative rates of roughly 70 to 83 per cent across schemes — meaning that under realistic forensic conditions, the mark was already missing most of the watermarked text it was looking at.
So this is not an ongoing contest. Removing a watermark costs one paraphrase pass and roughly zero effort. The tooling has already commercialised into a full product category with enterprise integrations and APIs, though the vendors’ “undetectable” claims are themselves overstated — a tool bypassing 90 per cent of detectors one month may manage 60 the next after an update. The race that exists is between removal tools and detectors, which is a different and much messier fight, and one where the collateral damage lands on people rather than on text.
What the mark actually measures now
Here is the consequence, and we think it is the whole story.
If removal is free and one step away, then the presence of a watermark tells you almost nothing about how a text was produced. It tells you about the person who produced it: that they did not know the mark was there, did not know it could be removed, or did not care enough to bother.
Every one of those correlates with not hiding anything. The person quietly passing off generated work as their own is precisely the person motivated to run it through a paraphraser. The student who used the assistant to fix their grammar and copied the result straight out is the one who keeps the mark.
The watermark, in other words, does not sort honest text from dishonest text. It sorts sophisticated users from naive ones. And it will be read as if it did the first thing.
The institutions closest to this are already discovering it. Independent testing of AI detectors has found false-positive rates around 11 per cent in some evaluations, and a study of real student submissions reported 18 per cent false positives alongside 32 per cent false negatives — flagging nearly one in five innocent students while missing a third of the actual cases. Yale, UCLA, Berkeley, UC San Diego, Waterloo, Michigan State and Vanderbilt have all disabled or restricted these tools. There have been lawsuits, at Yale in 2025 and Michigan this year.
That is what it looks like when a signal is treated as evidence after it has stopped being evidence.
Normalisation is coming anyway, and it inverts the market
Our reader’s first scenario is right, and its consequence is underrated.
When a corporate assistant rewriting an email produces marked text, and hundreds of millions of people use one, the mark becomes true of nearly everything. “AI-generated” stops carrying information for the same reason “typed rather than handwritten” stopped carrying information.
At that point the valuable signal inverts. The scarce claim is not “this was not machine-generated” but “this was verifiably human,” and that is expensive to establish — it needs process, provenance, attestation, someone’s reputation on the line. Anything both valuable and expensive to prove gets sold. We would expect a market in certified human authorship to emerge, and to be a better business than detection ever was, because it works with the grain of the incentives rather than against them.
The third path is the only stable one
Which brings us to the option we happen to argue for, and we want to be careful not to oversell it.
Running your own open-weight model does not win the arms race. It declines to enter it. You are not defeating a mark; you simply never generate one. That is structurally different from every removal tool, all of which depend on staying ahead of a detector that updates without warning.
The honest limits are real. It is available to a minority with the capability and a reason to bother. It does not remove disclosure obligations if you publish into a jurisdiction that requires them, as we noted last week. And it is not a way to be dishonest more effectively — if that is the appeal, the paraphraser was always cheaper.
What it does buy is the thing we keep returning to: the decision stays with you. Not the ability to hide, but the ability to choose what your systems emit, on your timeline, rather than finding out from a vendor announcement.
What a business should take from this
One practical instruction, and it is unambiguous: do not build a process that depends on detecting AI. Not in hiring, not in vendor quality assurance, not in academic or editorial integrity. The false-positive rates are high enough to generate accusations you cannot defend, and the false-negative rates are high enough that the people you are worried about will pass anyway. You would be buying the worst of both errors.
Build processes that do not need the distinction instead. Assess whether the work is correct, whether the reasoning holds, whether the claims check out. That is harder than running a detector, it is the thing you actually care about, and unlike provenance it cannot be removed with one paraphrase pass.