What a Detector Actually Sees: From Your Paragraph to a Probability, Step by Step
06 Jul 2026
You paste a paragraph, hit a button, and a number comes back: “72% AI.” It feels instantaneous, like the tool *read* your writing and formed an opinion. It didn’t. Between your paragraph and that percentage sits a short assembly line, and every station on it is doing something narrow and mechanical that has nothing to do with understanding what you wrote.
Walk the line once and the whole thing demystifies. You’ll see exactly where the number comes from — and, just as usefully, exactly why it can be so confidently wrong.
Key takeaways
- A detector never understands your meaning. It measures the statistical shape of your words.
- Your text passes through a pipeline: clean → tokenize → score each token → summarize into features → classify → squash into a percentage → aggregate.
- The core signal is per-token predictability, drawn from a scoring model’s probabilities and ranks.
- The final percentage is a sigmoid-squashed classifier output, not a direct reading.
- Knowing the steps makes the failures legible — the machine measures form, not authorship or truth.
Step 1: Cleaning the input
Before anything statistical happens, the detector tidies your text. It may strip out formatting and markup, normalize odd whitespace, and standardize characters. This step is quiet but consequential: if the tool strips markdown, those ## and ** symbols vanish before scoring; if it doesn’t, they ride along as unusual tokens. Either way, the version of your text that gets scored is often not byte-for-byte what you pasted. The cleaning stage already shapes the outcome, and you rarely get to see how it ran.
Step 2: Tokenization
Now the real work begins. The detector chops your cleaned text into *tokens* — words and word-fragments that the scoring model operates in. “Understanding” might become “under” + “standing”; a rare word might shatter into several pieces; a space is often part of a token too.
This matters more than it looks, because everything downstream is computed *per token*. The probabilities, the ranks, the final features — all of it is measured on this token sequence. That’s why oddities like look-alike Unicode characters or strange spacing can derail a detector: they change how the text tokenizes, and a word split into unfamiliar pieces gets unfamiliar probabilities before a single judgment is made. The tokenizer is the lens through which the model sees your writing, and a smudged lens distorts everything after it.
Step 3: Scoring each token
Here’s the heart of it. The detector runs your token sequence through a scoring model — usually not the model that wrote the text, but a stand-in — and for every token records how the model reacted. At each position the model had produced a full ranked list of possible next tokens with probabilities, and the detector reads off three kinds of information about the token you actually used:
- Probability (log-likelihood): how likely the model thought your token was. High means unsurprising.
- Rank: where your token sat in the model’s ranked list. Rank 1 means it was the top pick.
- Entropy: how spread out or confident the model’s whole prediction was at that spot.
If those terms are new, we walk through log-likelihood, rank, and entropy as token statistics in detail. The point here is that the detector now has, for every token in your paragraph, a small profile of how predictable it was. This is the raw material. Notice what it is *not*: it’s not meaning, not accuracy, not intent. Just predictability, token by token. The old GLTR tool made this literally visible by color-coding each token by its rank — a nice reminder that “what the detector sees” is a heatmap of predictability, nothing more.
Step 4: Summarizing into features
A paragraph has too many per-token numbers to feed a classifier directly, so the detector compresses them into a few summary statistics — the features the final decision is actually based on.
The two headline features are the familiar ones. Perplexity rolls up the token probabilities into a single measure of overall surprise; low perplexity means the text was predictable throughout. We work perplexity out by hand in what perplexity measures, worked out by hand. Burstiness measures how much that predictability *varied* across the piece — flat and even, or lumpy and human. Detectors often add more: average rank, entropy statistics, sentence-length variation, and so on.
Whatever the exact list, this is the moment your entire paragraph gets crushed down to a short vector of numbers. Everything you wrote, every idea and turn of phrase, is now a handful of features. The classifier will never see your words again — only this compressed summary.
Step 5: Classifying
Those features feed a decision function. In a trained detector it’s a classifier that learned, from labeled examples, which combinations of features tend to mean “human” and which mean “AI.” In a simpler zero-shot method it might be a threshold on a single statistic. Either way, the output at this stage is a raw, unbounded score — a *logit* — saying how far your feature summary sits toward the “AI” side of the boundary.
This raw score has no natural scale, so it’s not shown to you yet. It’s an internal number, meaningful only relative to the boundary the classifier drew.
Step 6: Squashing into a percentage
To make the raw score presentable, the detector runs it through a *sigmoid* — an S-shaped function that folds any number into the range between 0 and 1. Multiply by 100 and there’s your percentage.
This is where a confident-looking figure gets manufactured. Sigmoids flatten hard at the extremes, so once your feature summary is comfortably on one side of the boundary, the output rushes toward 0% or 100% and stays there. A “94% AI” can come from a raw score that’s only moderately past the line. The dramatic number is partly the math dramatizing itself, not the evidence being overwhelming.
Step 7: Aggregating across the document
For anything longer than a few sentences, the detector usually doesn’t score the whole thing as one lump. It breaks the text into chunks — often sentences — scores each one through the pipeline above, and then combines those into a document-level number, sometimes highlighting the individual sentences it found most suspicious.
That aggregation is why one weird, short, or predictable sentence can drag a whole document’s score around, and why the highlighted lines are usually your flattest, most predictable ones. The final percentage you see is a blend of many small per-chunk judgments, each with its own noise.
What the whole line adds up to
Trace the paragraph all the way through and the honest summary is stark. Your writing was cleaned, chopped into tokens, scored for predictability against a stand-in model, crushed into a few statistics, pushed across a learned boundary, squashed into a confident-looking percentage, and averaged across chunks. At no station did anything read your argument, weigh your evidence, or know whether a person or a machine sat at the keyboard.
That’s precisely why the pipeline can be so sure and so wrong at once. Plain, careful, or non-native writing produces low-perplexity, top-ranked tokens — statistically indistinguishable from machine text at every stage that matters — so the line flags it as AI while it was nothing of the sort. The failure isn’t a glitch in one step; it’s baked into a machine that measures *form* and reports it as if it were *authorship*. Seeing the steps is the antidote: a score is the exhaust of a predictability machine, and it deserves exactly the weight that description implies — a starting point for a closer look, never the last word.
Frequently asked questions
Does an AI detector actually read and understand my writing?
No. It never understands your meaning the way a person does. It converts your text into tokens, asks a model how predictable each token was, turns those numbers into statistics, and runs them through a classifier that outputs a probability. At no point does it grasp your argument or check your facts. It measures the statistical shape of the words, which is why it can flag perfectly true, perfectly human writing.
What is tokenization and why does it matter for detection?
Tokenization is the first real step: the detector breaks your text into tokens, the word-chunks the model works in. Everything downstream — probabilities, ranks, the final score — is computed per token, so how your text gets split shapes the result. It’s also why odd characters, unusual spacing, or look-alike Unicode can throw a detector off, since they change how the text tokenizes before any scoring happens.
What statistics does the detector compute from my text?
Usually a handful drawn from the model’s per-token predictions: perplexity (overall surprise), burstiness (how much that surprise varies), and often rank and entropy information about where your words fell in the ranked predictions. These few numbers are the compressed summary the final classifier judges — your whole paragraph boiled down to a short list of features.
How does the pipeline turn those numbers into a percentage?
The features feed a classifier or threshold that produces a raw score, and that score gets squashed by a sigmoid into a value between 0 and 1, shown as a percentage. For longer text the tool scores chunks separately and combines them into a document-level number. The percentage is the end of an assembly line — tokenize, score, summarize, classify, squash, aggregate — not a direct reading.
Why can the pipeline be confidently wrong?
Because every stage measures form, not truth. Plain, careful, or non-native writing produces low-perplexity, top-ranked tokens that look statistically machine-like, so the pipeline flags it despite being human. And the final sigmoid manufactures confident-looking numbers from modest evidence. Understanding the steps makes the failures legible: a confident score is the output of a form-measuring machine, not a verdict on authorship.
The short version
There’s no comprehension anywhere in an AI detector. Your paragraph gets cleaned, tokenized, scored for predictability against a stand-in model, summarized into perplexity and burstiness and a few friends, pushed across a learned boundary, squashed into a percentage by a sigmoid, and averaged across chunks. That assembly line measures the statistical form of your words and reports it as if it were authorship — which is exactly why it flags careful human writing with such confidence. Know the steps and the number loses its spell. Want to watch it happen to your own writing? Run a draft through the free checker and see which sentences light up, then browse more detection guides or see the plans.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.


