Watermarking vs Post-Hoc Detection: Two Fundamentally Different Strategies Compared
07 Jul 2026
There are two ways to answer the question “did a machine write this?” and they’re almost opposites. One plans ahead: mark the text at the moment it’s created, so you can check for the mark later. The other works backwards: take finished text you know nothing about and guess from its texture. These aren’t two flavors of the same tool. They’re two different bets about where the truth lives, one in a planted signal, the other in a statistical guess, and understanding the split explains most of the confusion around AI detection.
Key takeaways
- Watermarking plants a hidden statistical pattern in the model’s word choices *as it writes*. Detection is then a test for that specific pattern.
- Post-hoc detection analyzes finished text after the fact, guessing from surface statistics like perplexity whether it looks machine-made.
- Watermarking is reliable but narrow: it only works on text from a model that inserted the mark, and needs the model maker’s cooperation.
- Post-hoc detection is universal but shaky: it runs on any text, but it’s a guess, and guesses on human writing cause false positives.
- Both can be weakened by paraphrasing, at different rates. Neither is proof.
Two bets on where the truth lives
Start with the philosophical difference, because it drives everything practical.
A watermark is a signal someone deliberately put into the text. Before the model finishes a sentence, it subtly biases which words it picks, in a way that’s invisible to a reader but detectable by a matching statistical test. The truth about the text’s origin is *encoded in the text on purpose*. Verifying it later isn’t guessing; it’s checking for a fingerprint you know was planted.
Post-hoc detection makes no such assumption. It takes text that carries no planted signal and tries to infer origin from how the writing *behaves*, how predictable the words are, how much the sentences vary, how the token ranks fall. There’s nothing intentional to find, so the tool is reading tea leaves, sophisticated tea leaves, but tea leaves. The truth about origin isn’t in the text; the detector is inferring it from a proxy.
That’s the whole fork in the road. One method reads a message that was left for it. The other interprets writing that never had a message in it at all. Both are, in a sense, two big visions for tracing AI text, and they inherit completely different strengths and failures from that starting choice.
How watermarking works, briefly
The best-known approach comes from a 2023 paper by John Kirchenbauer and colleagues. The mechanism is elegant. At each step of generation, the model uses the previous tokens to seed a pseudo-random split of its vocabulary into a “green list” and a “red list,” then gently prefers green-list words. A human, and even a careful reader, notices nothing; the text still reads fine because there are always plenty of good green words to choose from.
But the statistics tell on it. Over a few hundred words, watermarked text contains far more green-list tokens than chance would ever produce. A verifier who knows the secret can run a statistical test, essentially “are there suspiciously many green words here?”, and get a clean answer with a very low false-positive rate. Newer production systems refine this idea considerably, but the principle holds: plant a pattern only a keyed test can see, then test for it. We walk through the mechanics in how statistical watermarking actually plants its signal.
The payoff is real: when a watermark applies, verification is far more trustworthy than any post-hoc guess. It’s checking for something it knows is there.
The catch that keeps watermarking narrow
So why isn’t detection a solved problem? Because a watermark can only be found in text where a watermark was *inserted*, and that requires the model maker to build it in and turn it on.
That leaves gaping holes. Open-source models anyone can download and run don’t watermark unless their operator chooses to, and someone trying to hide AI use certainly won’t. Older text predates any watermarking. And there’s no way to retroactively stamp a mark onto writing that was generated without one. So watermarking covers exactly one slice of the world: current text, from participating providers, that hasn’t had the mark stripped out. Everything outside that slice is invisible to it.
This is the fundamental trade. Watermarking buys reliability by narrowing its scope to text it had a hand in. It’s powerful inside the walled garden of a cooperating provider and completely blind beyond the wall. That narrowness is precisely why post-hoc detection persists despite being the less reliable method: it’s the only one that will even *attempt* a verdict on arbitrary text.
Post-hoc detection: universal reach, universal doubt
Post-hoc detectors, the perplexity checkers, the trained classifiers, the tools students and teachers actually use, have the opposite profile. Hand one any text and it will give you a score. That universality is genuinely useful; you don’t need anyone’s cooperation or secret key.
But every score is a guess about proxies, and the proxies overlap between humans and machines. Clear, conventional human writing looks statistically “predictable,” which is what these tools read as AI, so they false-flag careful and non-native English writers with grim regularity. And they’re easy to fool in the other direction: a light paraphrase or a few edits can push machine text out of range. OpenAI’s own post-hoc classifier caught only a fraction of AI text while flagging some human writing, and the company retired it in 2023 for low accuracy. Universal reach, but a verdict you can never fully trust.
Side by side
| Watermarking | Post-hoc detection | |
|---|---|---|
| When it acts | At generation, plants a signal | After the fact, guesses |
| What it checks | A pattern it knows was inserted | Surface statistics as a proxy |
| Reliability | High, low false positives, when it applies | Shaky, notable false positives |
| Coverage | Only watermarked models | Any text at all |
| Needs cooperation? | Yes, from the model maker | No |
| Beaten by | Heavy paraphrasing, stripping the mark | Light paraphrasing, edits, or just being a clean human writer |
Read the coverage row against the reliability row and you have the entire tension. The reliable method can’t reach most text; the method that reaches all text isn’t reliable. There is currently no strategy that is both universal and trustworthy, and that gap is the honest state of the field.
Where paraphrasing hits both
Neither strategy survives determined rewording, though they fail differently.
A watermark lives in a specific pattern of word choices, so replacing many of those words with a paraphraser disturbs the green-word surplus and erodes the signal. Good watermarks are designed to tolerate light editing, but robustness against aggressive, repeated paraphrasing is still an open research problem, one the 2023 analysis by Sadasivan and colleagues put squarely on the table. Post-hoc detectors are even softer targets; the same paper showed paraphrasing dragging several toward chance. So paraphrasing weakens both, just at different rates and for different reasons. We go deeper on the watermark side in whether watermarks survive paraphrasing and editing.
The takeaway isn’t “here’s how to beat them.” It’s that no tracing method, planted or guessed, is a magic certainty, and treating any single one as proof ignores a documented failure mode.
What this means for you
If you’re a writer, student, or content lead today, the method at your door is almost always post-hoc detection, and it’s the one prone to false-flagging writing you did yourself. That’s the crucial reframe: a post-hoc flag is a *guess about patterns*, not the discovery of a planted mark. It can be, and regularly is, wrong about honest human writing, because clean prose and machine prose share a statistical look.
Watermarking, by contrast, mostly matters if you’ve used a specific commercial model directly, and it’s checking for something that either is or isn’t there. Knowing which kind of claim you’re facing changes how much weight it deserves. A watermark hit is a strong signal about a specific model’s output; a post-hoc score is a probability that says more about your writing’s texture than its authorship. To see how the post-hoc tools read a piece before it becomes anyone’s problem, run a draft through the free checker and focus on the passages that read flattest.
Frequently asked questions
What’s the core difference between watermarking and post-hoc detection?
Timing and access. Watermarking plants a hidden statistical pattern in the model’s word choices as it writes, so a later test checks for that pattern. Post-hoc detection runs after the fact on text it didn’t create, guessing from surface statistics whether it looks machine-made. One reads a planted signal; the other guesses about text that carries none.
Which one is more reliable?
Watermarking, when it applies, because it verifies a signal it planted rather than guessing, with a low false-positive rate. But it only works on text from a model that inserted the mark. Post-hoc detection runs on any text but is a guess, and guesses on human writing cause false positives. Reliable-but-narrow versus universal-but-shaky.
Why doesn’t everyone just use watermarking?
Because it needs the model maker’s cooperation and only marks that model’s output. It can’t appear retroactively in text from a model that never inserted it, including open-source models. So it covers only the slice of AI text from participating providers and misses everything else, which is why post-hoc detection still exists.
Can paraphrasing remove a watermark?
It can weaken one. A watermark lives in a pattern of word choices, so heavy rewording disturbs it, especially aggressive or repeated paraphrasing. Good watermarks survive light editing, but robustness against determined paraphrasing is an open problem. Post-hoc detectors are even easier to fool this way, so neither is immune.
As a writer, which one affects me?
Mostly post-hoc detection, since that’s what schools and platforms run, and it’s the one that false-flags honest writing. Watermarking mainly affects text from specific commercial models. A post-hoc flag on your own writing is a guess about patterns, not detection of a planted mark, and it can be wrong.
The bottom line
Watermarking and post-hoc detection are two answers to the same question that share almost nothing under the hood. One plants a signal and later tests for it, earning reliability at the cost of only ever seeing text it marked. The other guesses from the texture of arbitrary writing, earning universal reach at the cost of guessing wrong on honest humans. Neither survives determined paraphrasing, and neither is proof. Whichever you’re facing, the sane posture is the same: treat the result as a signal about a specific kind of pattern, weigh it against everything else you know, and never let a single number stand in for judgment. To see how the tools you’re most likely to meet read your writing, run a draft through the free checker, browse more on how detection works, or see the plans.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.


