The Difference Between a Detector’s Reported Score and Its Internal Logit
07 Jul 2026
When a detector tells you a passage is “97% AI,” it feels like a probability, a measurement, something you could bet on. It isn’t, quite. That tidy percentage is the last stop on an assembly line, and the raw material it started as, a number called a logit, has been squashed and reshaped so thoroughly that a lot of its meaning got left on the floor.
This is one of those under-the-hood details that sounds academic until you realize how much false confidence it manufactures. The gap between what the model computed and what the screen shows you is exactly where people go wrong.
Key takeaways
- Inside, a classifier computes a logit: a raw, unbounded number whose size and sign roughly encode how far the text sits from the decision boundary.
- A sigmoid function squashes that logit into the 0-to-100% you see. The squash is steep in the middle and flat at the ends.
- Because of that shape, the percentage compresses the confidence information. Two “99%” results can hide very different logits.
- A reported percentage is only a true probability if the tool is well calibrated, and most consumer tools aren’t audited for it.
- You almost never get to see the logit, the threshold, or the calibration. You get the polished number and are asked to trust it.
Meet the logit
Every neural classifier, an AI detector included, ends in a step that produces a raw score before any probability exists. For a two-way decision like AI-versus-human, you can think of it as a single number, the logit, that can be anything: strongly negative, zero, strongly positive.
The sign tells you which way the model is leaning. The magnitude tells you how hard. A logit of +8 means “way over on the AI side, not close.” A logit of -6 means “clearly human, not close.” A logit of +0.2 means “leaning AI, but barely, this is close to a coin flip.” That last case is the important one: the logit carries *distance from the boundary*, which is the closest thing the model has to an honest confidence.
If you could read the logit directly, you’d learn something the percentage hides: not just which side, but how emphatically. A passage scoring +0.3 and a passage scoring +9 are worlds apart in the model’s internal state, even if they end up wearing the same label.
The sigmoid squash, and what it destroys
The logit is unbounded, and nobody wants to see a score of “+8.4.” So detectors run it through a sigmoid, an S-shaped function that maps any real number into the 0-to-1 range. Big positive logits get mapped near 1 (shown as ~100%), big negative logits near 0, and a logit of zero maps to exactly 0.5, the fence.
Here’s the mischief. The sigmoid is steep in the middle and nearly flat at the edges. Around a logit of zero, a small change produces a big move in the percentage. Out at a logit of +6 or +8, an enormous change in the logit barely nudges the percentage at all, because it’s already pinned near 100%.
So the percentage does two unhelpful things at once. Near the boundary, it *exaggerates*: a logit drifting from +0.1 to +0.5, still deeply uncertain, might show as a jump from 52% to 62%, which reads like a real difference and isn’t. And at the extremes, it *compresses*: a logit of +4 and a logit of +12 both display as “99%,” even though the second is vastly more confident than the first. The number you see throws away the very thing you’d most want to know, how far past the line the model actually is.
A mini-scenario: two identical “99%” verdicts
A teacher runs two essays. Both come back “99% AI.” Feels like an open-and-shut match, two cases of the same thing.
Under the hood, essay one scored a logit of +3.1, just far enough past the boundary to pin near 100% on the sigmoid. Essay two scored +11. Essay two is machine-flat in a way that’s statistically unmistakable. Essay one is only modestly over the line, the kind of result a paraphrase or a length change could have flipped. The interface flattened both to the same “99%,” erasing a difference that should absolutely change how the teacher treats them. The percentage lied by omission, not by being wrong, but by hiding the confidence gap the logits made obvious.
That’s the practical cost of the squash: it makes a shaky result and a rock-solid result look identical.
Is the percentage even a real probability?
Now the deeper question. Even setting aside the compression, is “90%” a claim that there’s a 90% chance the text is AI?
Only if the detector is calibrated. Calibration is the property that, across all the texts a tool scores 90%, about 90% of them genuinely are AI. It’s a real, measurable thing, and it does not come for free. A now-classic 2017 study by Guo and colleagues showed that modern neural networks are frequently *miscalibrated*, often overconfident, with their reported probabilities drifting away from true frequencies. Detectors inherit that problem. A tool can be a decent classifier and still have percentages that don’t mean what they say. We dig into what calibration actually means for a detector on its own, because it’s the hinge the whole “is this a probability?” question turns on.
And there’s more processing on top. Many tools don’t show the raw sigmoid output at all. They apply their own scaling, bucket results into “likely/possibly/unlikely,” or slide the reported number relative to a decision threshold the score gets compared against that isn’t at 50%. By the time a percentage reaches your screen, it may have been squashed, calibrated (or not), rescaled, and bucketed. Reading it as a literal probability is reading precision that was never there. OpenAI, when it published and later retired its own classifier, reported results in coarse likelihood bands rather than a crisp percentage, an implicit admission that a sharp number would overstate what the tool knew.
Why the opacity is the real issue
Step back and the pattern is uncomfortable. The model has a genuinely informative internal quantity, the logit, that says how confident it is and how far from the line. Then every layer between that number and you *reduces* the information: the sigmoid compresses it, calibration may distort it, display choices bucket it. And at the end, you’re shown one clean percentage and invited to treat it as fact.
You can’t see the logit. You can’t see the threshold. You usually can’t see whether the tool was ever calibrated. Consumer detectors hand you the polished output and hide the machinery that would tell you how much to trust it. Some research APIs expose raw scores, but the tools students, teachers, and writers actually meet do not.
That’s exactly why an easily-flipped or borderline result should never carry the weight people give it. A “99%” might be a +11 logit or a +3, and you’ll never be told which. To get a feel for how a real draft scores, and to compare pieces rather than fixate on one decimal, you can run a draft through the free checker and watch which passages actually shift.
Frequently asked questions
What is a logit in an AI detector?
It’s the raw, unbounded number the classifier computes before any probability exists. Its sign says which way the model leans, its magnitude says how far the text sits from the decision boundary. Near zero means genuinely unsure. The logit is the model’s honest internal state; the percentage is that logit after squashing and display processing.
How does a logit become the percentage I see?
A sigmoid maps the unbounded logit into a 0-to-1 range, shown as a percentage. But the sigmoid is steep in the middle and flat at the ends, so small logit changes near the boundary swing the percentage a lot, while big changes at the extremes barely move it. The percentage compresses the distance-from-boundary information you’d most want.
Is a reported 90% the same as a 90% chance the text is AI?
Only if the tool is well calibrated, meaning 90%-scored texts really are AI about 90% of the time. Most tools aren’t audited for this and add their own scaling and bucketing, so read “90%” as “leans fairly strongly AI,” not as a literal nine-in-ten probability.
Why does this difference matter for me?
Because people treat the percentage as a hard probability when it usually isn’t. Two “99%” results can hide very different logits, one just over the line, one far past it, and you’d never know. Knowing the number is processed and compressed keeps you from reading false precision into it.
Can I see the raw logit?
Almost never. Consumer detectors show the polished percentage and hide the logit, threshold, and calibration. Some research tools expose raw scores, but the ones writers and students actually use don’t, which is part of why a single number deserves skepticism.
The bottom line
The percentage on a detector’s dashboard is not the number the model computed. It’s a logit that got squashed by a sigmoid, maybe calibrated, maybe not, then rescaled and bucketed for display, losing most of its confidence information along the way. Two identical “99%” results can sit on wildly different internal footing, and a “90%” is a real probability only if someone bothered to make it one. Read the reported score as a leaning, not a verdict, and distrust its precision, because the machinery that would justify that precision is exactly what you’re not allowed to see. To compare how your own drafts score without over-reading one decimal, run a draft through the free checker, browse more on how detection works, or see the plans.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.


