What ‘AI Probability’ Numbers Hide: Sigmoid Outputs vs Real-World Odds
06 Jul 2026
A detector hands you “87% AI” and your stomach drops. It sounds like a probability, phrased like one, colored red like a warning. So the natural reading is: there’s an 87% chance a machine wrote this. That reading is wrong, and the gap between what the number *feels* like and what it *is* has gotten real students in real trouble.
Let’s take the number apart. Once you see where it comes from, you’ll stop reading it as a verdict and start reading it as what it is — one model’s confidence, on a scale it made up.
Key takeaways
- The percentage isn’t a probability of authorship. It’s the detector’s internal confidence, reshaped to look like one.
- That reshaping is done by a sigmoid function, which squashes any raw score into the 0-to-1 range no matter how meaningful the underlying score is.
- Sigmoids saturate, so confident-looking numbers like 99% can come from fairly modest evidence.
- A well-calibrated detector is rare. Most neural classifiers are overconfident unless someone deliberately fixes it.
- Even a good percentage is worthless without the base rate — how common AI text is in the batch being checked.
Where the percentage comes from
Under the hood, most trained detectors don’t compute a probability directly. They compute a raw, unbounded score — the jargon is a *logit* — that says how far your text sits on the “AI” side of a decision boundary. A big positive logit means “deep in AI territory.” A big negative one means “clearly human.” Zero sits right on the fence.
That raw score is ugly to show a user. It has no natural ceiling, no floor, no intuitive scale. So the detector runs it through a *sigmoid*, an S-shaped function that takes any number from negative infinity to positive infinity and politely folds it into the range between 0 and 1. Multiply by 100 and you’ve got your percentage. Clean, familiar, and — this is the trap — indistinguishable on the screen from a genuine probability.
The sigmoid makes it look like a probability
Here’s the sleight of hand, and it isn’t malicious, just mathematical. A sigmoid’s output is *always* a number between 0 and 1. That’s its entire job. So any classifier, calibrated or not, accurate or not, will emit something that reads like a probability. The format is guaranteed. The meaning is not.
Think of it like a thermometer that always displays a value between 0 and 100 degrees, but whose calibration nobody ever checked. It’ll confidently tell you it’s 87 degrees. Whether that corresponds to anything real about the temperature is a separate question the display can’t answer. The sigmoid is that display. It promises you a nicely bounded number and says nothing about whether the number is true.
Why 99% is less impressive than it looks
Sigmoids have a specific shape that matters for reading these scores: they saturate at the ends. In the middle, small changes in the raw score move the percentage a lot. Out at the extremes, the curve flattens hard, so once your text is *somewhat* past the boundary, the output rushes toward 100% and stays there.
The practical effect is that extreme percentages are cheap. A “99% AI” verdict doesn’t require the model to be nearly certain in any grounded sense — it just requires a raw score comfortably on one side of the line, and the sigmoid does the dramatizing. That’s part of why detector outputs feel so much more confident than the underlying evidence warrants. The math manufactures certainty at the tails.
Calibration: the step almost nobody does
There’s a real version of “87% means 87%.” It’s called calibration, and a calibrated detector is one where, of all the texts it labels “80% AI,” genuinely about 80% turn out to be AI-written. That’s the property you actually want when a number is going to be used as evidence.
Most classifiers don’t have it. A well-known 2017 paper by Guo and colleagues, *On Calibration of Modern Neural Networks*, showed that modern deep networks are systematically overconfident — their reported probabilities run hotter than their real accuracy. Fixing this takes deliberate work, like *temperature scaling*, where you divide the logits by a tuned constant before the sigmoid to cool the overconfidence down. Unless a detector has done that calibration and validated it on text like yours, its percentages are ranked confidence — useful for sorting, not for betting. This is closely related to why a single flagged line can distort a whole document’s score, which we get into in why one flagged sentence can sink a whole essay.
The base rate the number can’t see
Even a perfectly calibrated detector hides something the percentage on your screen literally cannot include: how common AI text is in the first place. This is the base rate, and ignoring it is a famous reasoning error for a reason.
Run the thought experiment. Suppose you’re a teacher with 100 essays, 95 of them honestly human-written, and you run a detector that’s 90% accurate in both directions. On the 5 real AI essays it flags about 4 or 5. But on the 95 human essays it also *wrongly* flags about 9 or 10. So of everything it marks “AI,” more than half is human work. A flag, in that setting, is more likely a false alarm than a catch — despite a detector that sounds quite accurate. The percentage never told you this, because it can’t. It doesn’t know the mix of the pile it’s drawing from.
Flip the base rate and the story flips too. In a batch that’s mostly AI-generated spam, the same flag becomes far more trustworthy. The number on the screen is identical; its real meaning is completely different. You have to supply the context yourself.
Reading the number honestly
None of this makes detectors useless. It makes them instruments you have to read correctly. A high score is a reason to *look closer* — at drafts, revision history, sources, the writer’s own account — not a reason to conclude. The percentage is a starting point for a conversation, not the end of one.
It also reframes what the different detection approaches are even doing under the hood. Whether a tool is a trained classifier or a zero-shot method, the final flourish is usually the same sigmoid squeeze, and the same caveats apply. When someone waves a percentage at you as proof, the right questions are: calibrated against what data, and against what base rate? Most of the time, nobody can answer either.
Frequently asked questions
Does ‘87% AI’ mean there’s an 87% chance a machine wrote my text?
No. The 87% is the detector’s internal confidence, squeezed into the 0-to-100 range by a sigmoid function. It reflects how far your text sits on the “AI” side of the model’s boundary, not a calibrated probability of authorship. Whether it maps to a real 87% depends on calibration and on how common AI text is in the batch being checked, and usually it doesn’t map cleanly at all.
What is a sigmoid and why does it matter here?
A sigmoid is an S-shaped function that takes any number and squashes it into a value between 0 and 1, shown as a percentage. Detectors compute a raw “logit” and pass it through the sigmoid so the output looks like a probability. The catch is that looking like a probability and being a well-calibrated one are different things. The sigmoid guarantees the first, not the second.
Why do detectors show such confident-looking numbers like 99%?
Because the sigmoid saturates. Once your text lands far enough past the boundary, the function flattens toward 1, so a moderately confident model produces a very extreme-looking percentage. A 99% can come from a raw score that’s only somewhat over the line. The dramatic number is partly an artifact of the math.
What does calibration mean for an AI detector?
A perfectly calibrated detector would be right about 80% of the time on everything it labels “80% AI.” Most classifiers aren’t calibrated out of the box — modern networks tend to be overconfident — and methods like temperature scaling exist to fix that. Unless a tool has done that work and validated it on data like yours, its percentages are ranked confidence, not trustworthy odds.
Why does the base rate change how I should read a score?
Because a test’s real-world reliability depends on how common the thing is. If almost everything in a batch is human-written, even a fairly accurate detector produces false positives that outnumber true positives, so a flagged piece is often still more likely human. The percentage ignores that context entirely — you have to supply the base rate yourself.
The short version
The “87% AI” on your screen is a sigmoid output dressed as a probability. It tells you how confident one model is, on an uncalibrated scale, with the base rate stripped out — three quiet omissions that together mean the number is a hint, not a fact. Extreme scores are especially easy to over-read, since the sigmoid manufactures certainty at the tails. Read detector percentages the way you’d read a smoke alarm: worth investigating, never proof on its own. Want to see how your writing scores and, more usefully, *why*? Test a draft in the free checker, then read more on how detection works or compare the plans.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.


