How Detectors Estimate Perplexity Without Access to the Original Model
06 Jul 2026
Here’s a problem most explainers skip right past. Perplexity is supposed to measure how predictable your text was — but predictable *to which model?* A detector staring at your essay has no idea what produced it. Maybe GPT-4. Maybe Claude. Maybe Gemini, maybe an open model someone fine-tuned in a garage, maybe a person. It can’t call up the author and borrow their probabilities.
So how does it put a perplexity number on your writing at all? The short answer is that it cheats, in a principled way. It brings its own model to the table.
Key takeaways
- Perplexity is model-relative. There’s no such thing as the perplexity of a text on its own — only its perplexity according to some specific model.
- Detectors don’t have the model that wrote your text, so they score it with a stand-in, usually called a proxy or surrogate model.
- This works because predictability is largely shared. What one model finds obvious, most others find obvious too.
- The catch is mismatch. When the proxy and the true author disagree, the signal blurs, which is part of why different tools give you different scores.
- The number is a real measurement of one real thing, and still not a verdict.
Perplexity needs a model to exist
Start with the thing that trips people up. When a tool reports that your text has “low perplexity,” it sounds like an objective property, the way mass or length is. It isn’t. Perplexity is a reading taken *by* a model, and it only means anything relative to that model’s expectations. If you want the ground-level version of what the number measures, we walk through what perplexity actually measures in a separate piece.
The consequence is subtle but important: the same sentence has a different perplexity for GPT-2 than for GPT-4, because those two models learned slightly different senses of what’s likely to come next. So a detector faces an awkward truth. To score your text, it needs *a* model. It just doesn’t need — and can’t have — the *right* one.
Enter the proxy model
The workaround is to pick a model, any reasonable model, and use it as a yardstick. This stand-in goes by a few names: proxy model, surrogate model, reference model, scoring model. They all mean the same thing — the model the detector actually loads and runs your text through to get its probability numbers.
You might expect detectors to reach for the biggest, newest model they can find. Often they do the opposite. Plenty of research-grade detectors score text with a comparatively small, older, open model — a GPT-2 variant is a classic choice — for very practical reasons. Small models are cheap to run on every submission, they’re openly available (you can’t easily get raw token probabilities out of most commercial APIs), and, surprisingly, they don’t seem to lose much accuracy for the trouble.
Why a smaller stand-in can still work
This is the part that feels like it shouldn’t hold up, and yet it does. How can a modest 2019-era model judge text written by a 2026 giant?
Because predictability travels. Language models trained on overlapping piles of internet text end up agreeing, to a large degree, on what counts as an expected word. “The results were statistically significant” is a low-surprise phrase for basically every model that has ever read a research paper. You don’t need the exact author model to notice that a passage took the most expected path at every single step — that flatness shows up almost regardless of which model is holding the ruler.
There’s a genuinely counterintuitive research finding here. A 2023 paper by Mireshghallah and colleagues, titled *Smaller Language Models are Better Zero-Shot Machine-Generated Text Detectors*, reported that small models can outperform large ones at spotting machine text. Part of the intuition is that a smaller model is more easily “surprised,” so the gap between its reaction to flat AI prose and to lumpier human writing can actually be wider. The undersized yardstick isn’t a compromise so much as a decent tool in its own right.
DetectGPT, the curvature method from Mitchell and colleagues at Stanford, leans on the same reality from a different angle — it queries a scoring model’s probabilities without ever needing to be the model that generated the text. The whole zero-shot detection genre rests on this transfer of predictability from one model to another.
Where the trick quietly breaks
None of this is free. The proxy is an approximation, and every approximation has a failure mode.
Mismatch blurs the signal. When the proxy model and the true author model disagree about what’s predictable, the measurement gets noisy. Text from a very new, very unusual, or heavily fine-tuned model can behave in ways the humble proxy didn’t anticipate, and the perplexity reading drifts away from what you’d get with the real author in hand. The best signal, as the DetectGPT authors themselves note, comes when you score text with the exact model that wrote it — a condition that almost never holds in the wild.
Different tools, different proxies. Every detector picks its own scoring model, and there’s no rule that two proxies must agree. That’s a big reason the same paragraph reads as “98% human” in one tool and “AI-generated” in the next. You’re not seeing one truth reported two ways; you’re seeing two different rulers.
Humans who write plainly get caught in the middle. The proxy can’t distinguish “predictable because a machine wrote it” from “predictable because a careful person chose clear words.” That limitation isn’t specific to proxies, but proxies inherit it fully. It’s the same mechanism behind documented false positives against non-native English writers, whose more conventional phrasing scores low no matter whose model does the scoring.
You can’t see the ruler. Commercial detectors rarely disclose their scoring model, and some fold perplexity into a stew of other signals. So you can’t calibrate against their scale, and the scale itself can shift the day they swap in a new proxy. Any number you memorize is written in disappearing ink.
What this means for your writing
The practical upshot is oddly reassuring. Since you can’t know or control which proxy a given tool uses, there’s no clever number to chase. The only durable move is to make your writing genuinely less predictable to *any* reasonable model — and the way you do that is the same way you make it better for human readers.
Reach for the specific detail instead of the generic one. Let sentence lengths vary instead of settling into a hum. Say the slightly surprising thing you actually mean rather than the smoothest available cliché. Every one of those choices raises perplexity against every proxy at once, because it makes your next word harder for *any* model to guess. You’re not gaming a particular ruler; you’re stepping off the flat, over-expected path that all the rulers are built to catch. If you want to see which of your lines sit on that path, the token statistics detectors read are worth understanding too.
Frequently asked questions
If a detector doesn’t have the model that wrote my text, how does it measure perplexity at all?
It uses a stand-in. Perplexity is always relative to some model, so the detector loads its own scoring model — often a smaller open one like a GPT-2 variant — and asks how surprised it is by your words. The number isn’t “the” perplexity of your text. It’s the perplexity according to the detector’s proxy, which may be a very different model from whatever wrote the passage.
Why would a smaller, older model be any good at judging text from a newer one?
Because predictability is largely shared across models. The phrasings a big model finds obvious tend to be obvious to a small one too, since they learned from overlapping data and the same regularities of English. Research has even found smaller models can be better zero-shot detectors. You don’t need the exact author model to notice a sentence took the most expected path at every step.
Does using a proxy model make detection less accurate?
It adds a source of error. When the proxy and the real author model disagree about what’s predictable, the signal blurs, and text from an unusual or very new model can score in unexpected ways. This mismatch is part of why the same passage reads as “human” in one tool and “AI” in another — they’re scoring with different proxies that were never guaranteed to agree.
Can I tell which model a detector uses to score my writing?
Almost never. Commercial tools rarely disclose their scoring model, and some blend several signals so there isn’t one clean answer. That opacity is why chasing a target perplexity number is a losing game — you can’t calibrate against a scale you can’t see, and it shifts whenever the tool swaps its proxy.
So is a perplexity-based score meaningless?
Not meaningless, just partial. It’s a real measurement of one real thing: how predictable your text looked to one particular model. That’s useful as a hint. It stops being useful the moment you treat it as a verdict, because the proxy could be mismatched to the true author, and predictable writing is something plenty of humans do on purpose.
The short version
A detector can’t borrow the model that wrote your text, so it brings its own — a proxy model, often smaller and older than you’d guess — and measures perplexity against that. It works better than it has any right to, because models mostly agree on what’s predictable. It also has real seams: mismatched proxies, tool-to-tool disagreement, and the age-old problem of plain human writing scoring low. Treat the number as one model’s opinion, not a fact about your words. Curious what your own draft looks like to a scoring model? Run a draft through the free checker and watch which lines light up, then browse the rest of our detection guides or see the plans.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.


