How Detectors Aggregate Per-Sentence Scores Into a Single Document Verdict
07 Jul 2026
A detector hands you one number for a whole essay, but it didn’t compute one number. It computed dozens, one per sentence, and then squeezed them into the single verdict you see. The squeezing step gets almost no attention, which is strange, because it can matter more than the sentence scores themselves. The same set of sentences can produce “12% AI” or “88% AI” depending on nothing but how the tool chose to combine them.
That combining rule is invisible, unstandardized, and quietly decisive. It’s worth understanding, because it’s where a lot of unfair verdicts are actually made.
Key takeaways
- Detectors score each sentence for how machine-like it looks, then apply an aggregation rule to produce one document score.
- Common rules: average the sentence scores, take the maximum (worst sentence), or count the fraction of sentences over a per-sentence threshold.
- The rule changes the verdict on identical text. Max-style rules let a single flat sentence condemn a whole human document.
- Aggregation is rarely disclosed, so two tools can disagree on the same essay purely because they combine differently.
- Sentence highlighting is a good editing map but not proof; a highlighted sentence is a low-confidence guess.
Detection is a two-step process
Most people picture a detector reading an essay and forming an impression of the whole. Under the hood it’s more mechanical, and it happens in two distinct stages.
Stage one: score the parts. The tool breaks your document into units, usually sentences, and scores each one for how predictable or machine-like it looks on the same statistics it always uses, perplexity, burstiness, token ranks. Now it has a list: sentence 1 scored low, sentence 2 scored high, sentence 3 medium, and so on down the page. This is the granular view some tools expose as colored sentence highlighting, and it’s the level GPTZero describes working at when it explains how its detector highlights sentence by sentence.
Stage two: combine into a verdict. Nobody wants a list of forty numbers, so the tool collapses them into one. This is aggregation, and it’s a separate decision from how the sentences were scored. Two detectors could score every sentence identically and still hand you different document verdicts, because they collapse the list differently. If you want the framing in full, we compare the difference between sentence-level and document-level scoring directly.
The rest of this piece is about stage two, because it’s the one nobody talks about and the one that quietly decides borderline cases.
The three ways to collapse a list
There are really only a few sensible ways to turn many sentence scores into one, and each carries a different personality.
Average. Add up the sentence scores, divide by the count. This is the democratic option: every sentence gets an equal vote, and no single one dominates. Averaging is *forgiving*. A document that’s mostly human with a few AI-looking sentences will still average out low. The downside is the mirror image: a genuinely AI document with a few varied sentences mixed in can average its way under the line. Averaging protects the innocent and lets some guilty slide.
Maximum (or near-max). Take the highest, most AI-looking sentence and let it stand in for the document, or weight heavily toward the worst offenders. This is the prosecutorial option. Its logic: if even one sentence screams machine, that’s worth flagging. The problem is severe. A single flat, conventional sentence, a plain definition, a stated fact, a line of boilerplate, can drag an entire human-written essay over the threshold. Max favors catching AI at the direct expense of honest writers.
Proportion over threshold. Count how many sentences individually cross a per-sentence cutoff, and base the verdict on that fraction, “60% of sentences read as AI.” This sits between the other two. It’s less hostage to one outlier than max, but a document sprinkled with short, conventional sentences can still rack up a high count, since short and plain sentences are exactly the ones that false-flag.
None of these is neutral. Each one is a policy about which mistake to tolerate, dressed up as a math operation.
A worked example: the same essay, three verdicts
Take a real-feeling case. A student writes a 20-sentence history essay. Eighteen sentences are lively and specific, scoring low (human). Two are flat by necessity: a textbook definition of “mercantilism” and a plainly-stated date-and-place fact. Those two score high (machine-like), not because a machine wrote them, but because there’s really only one clear way to state a definition, and clarity reads as predictable.
- Average: eighteen low scores and two high ones average out well under the line. Verdict: human. Fair result.
- Maximum: the document inherits the score of its single most AI-looking sentence, the definition. Verdict: likely AI. A false accusation built entirely on one unavoidable sentence.
- Proportion: two of twenty sentences flagged, 10%. Depending on the cutoff, probably a pass, maybe a borderline. Verdict: leans human, with a caveat.
Same eighteen good sentences, same two flat ones, three different fates. The student did nothing different. The tool’s aggregation rule wrote the verdict, and the student never got to see which rule was used or why. This is a big part of why detectors disagree with each other on identical documents, and why how GPTZero computes its score from sentence highlighting can land differently from a competitor scoring the same text.
Why max-style aggregation is so dangerous
The averaging-versus-max choice deserves a hard look, because max-style rules cause a specific, predictable injustice.
Almost every real document contains a few sentences that *must* be flat. Definitions have to be plain. Cited facts have to be stated straight. Transitions and topic sentences are often conventional by design. Instructions, methods sections, and legal boilerplate are flat on purpose, because flatness is clarity there. These aren’t signs of a machine; they’re signs of writing that’s doing its job.
A max-style aggregator treats every one of those necessary flat sentences as a potential smoking gun. It only takes one to tip the verdict. So the documents most exposed to false accusation under max aggregation are, perversely, the *careful* ones, technical writing, academic writing, and the conventional prose of non-native English speakers. The 2023 Stanford study in *Patterns* that found detectors misclassifying most non-native English essays is partly a story about this: conventional, low-variance sentences reading as machine-like, and aggregation rules that punish even a handful of them. A school that wants to avoid false accusations should specifically want a tool that doesn’t hang a verdict on its single worst sentence.
What per-sentence scores are actually good for
Here’s the constructive turn. Sentence-level scoring is genuinely useful, just not as evidence. It’s useful as a *map*.
When a detector highlights certain sentences as machine-like, set aside the question of whether AI was involved and read the highlight literally: these are the flattest, most predictable, most uniform sentences in your draft. That’s a real and actionable signal for a writer. Those are the lines worth revisiting, the places to vary a sentence length, add a concrete specific, or replace a generic phrasing with the exact one your meaning wanted. Used this way, the highlighting improves your writing whether or not a machine ever touched it. To get that map for a piece you’re working on, you can run a draft through the free checker and treat the flagged sentences as an editing to-do list rather than a confession.
What per-sentence scores are *not* good for is proving authorship of any single line. A highlighted sentence is a low-confidence guess about a short, noisy span, exactly the conditions where detection is least reliable. Even OpenAI retired its own classifier in 2023 over low accuracy, and short-span judgments are among the shakiest a detector makes. Read the map; don’t hang anyone with it.
Frequently asked questions
How does a detector turn sentence scores into one document verdict?
It scores each sentence, then applies an aggregation rule: average all the scores, take the highest, or count how many cross a per-sentence threshold. Each rule yields a different document verdict from the same sentences, because each weighs outliers differently, and tools rarely disclose which rule they use.
Why can one sentence flag a whole document?
Because some tools aggregate by the maximum, letting the document score reflect its most AI-looking sentence. Under that rule, a single flat sentence, a definition, boilerplate, a plainly stated fact, can drag an entire human document over the line. It’s a design choice favoring catching AI over protecting honest writing.
Is averaging or max better?
Neither in the abstract; they encode different priorities. Averaging is forgiving and can miss real AI use; max is aggressive and condemns honest documents with a flat sentence or two. Averaging favors human innocence, max favors catching AI. Schools worried about false accusations should want a tool that doesn’t rely on a single worst sentence.
Does sentence highlighting mean the tool is more accurate?
Not necessarily. Highlighting feels granular but each sentence score inherits the usual problems, short spans are noisy, conventional sentences look machine-like, and the aggregation on top can distort things. It’s a useful diagnostic for finding flat sentences, but a highlighted line is a low-confidence guess, not proof.
How should I use per-sentence scores in my own editing?
As a map, not a verdict. Highlighted sentences are the flattest, most predictable spots in your draft, worth adding variety or specificity to regardless of whether AI was involved. That’s genuinely useful. Just don’t read a highlight as proof; it’s pointing at uniform writing, which humans produce constantly.
The short version
A document verdict is a summary of many sentence scores, and the summary rule, average, max, or proportion, does as much to decide the outcome as the sentences do. Max-style rules are the quiet villains: they let one unavoidable flat sentence condemn an honest document, and they hit careful and non-native writers hardest. You almost never get told which rule ran, which is a big reason detectors disagree. Use sentence highlighting as an editing map to find and fix your flattest lines, and refuse to treat any single highlighted sentence as proof of who wrote it. To get that map for your own draft, run it through the free checker, browse more on how detection works, or see the plans.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.


