Precision vs Recall in AI Detection: Which One Should a School Actually Optimize?
07 Jul 2026
Every AI detector is secretly making a choice on your behalf, and most schools never notice it. The choice is which kind of mistake to make more often: letting a cheater slip through, or accusing an honest student. You can’t minimize both at the same time. The tool, or whoever configured it, already picked. The only real question is whether they picked the way your institution actually wants.
That’s the whole tension behind two words that sound like jargon but describe something very human: precision and recall.
Key takeaways
- Precision: of everything flagged as AI, how much really was? Low precision means false accusations.
- Recall: of everything that really was AI, how much got caught? Low recall means cheaters slip through.
- They trade against each other through one dial, the decision threshold. Pushing one up pushes the other down.
- For schools, the two mistakes have wildly different costs, so the honest answer is to optimize for precision, even though it lets some AI use go uncaught.
- The base rate makes this sharper: when few students actually cheat, even a small false-positive rate can drown the true catches.
Two questions, not one
People say a detector is “accurate” as if accuracy were a single number. It isn’t, not in any way that helps you. Split it into two questions and the whole picture changes.
Imagine a class of 100 essays. Ten students genuinely used AI. A detector flags 15 essays total.
Precision asks about those 15 flags. How many of the flagged essays were actually AI? If 9 of the 15 were real AI and 6 were honest students caught by mistake, precision is 9 out of 15, or 60%. Precision is the reliability of a positive result, the answer to “if I got flagged, how worried should I be?”
Recall asks about the 10 real AI essays. How many did the tool catch? If it got 9 of the 10, recall is 90%. Recall is the completeness of the sweep, the answer to “how much did we miss?”
Same detector, same class, two totally different numbers describing two totally different worries. And crucially, you can push one up by letting the other fall.
Why you can’t have both
Under the hood, a detector produces a score, and somewhere there’s a cutoff. Text above the cutoff gets flagged. That cutoff, the decision threshold, is the single dial that controls the whole tradeoff.
Slide the threshold down so the tool flags more aggressively. Now it catches nearly every real AI essay (recall climbs) but it also snares more borderline human writing (precision drops). Slide it up so the tool only flags the most blatant cases. Now almost everything it flags is genuinely AI (precision climbs) but a lot of lighter AI use sails through (recall drops).
There’s no magic setting where both hit 100% on real, messy student writing. This isn’t a temporary engineering shortfall; it’s baked into the geometry of a single threshold dividing two overlapping piles of text. Human and AI writing overlap in the middle, and no line through that overlap separates them cleanly. Every vendor picks a point on that curve. Most don’t tell you where, and the default is rarely tuned for a school’s actual risk.
The costs aren’t symmetric
Here’s the part that decides the whole argument. In a classroom, the two mistakes do not cost the same.
A false negative, missing a student who used AI, is a missed enforcement. It’s a real cost. That student got an unearned advantage this once. But the sky doesn’t fall. There will be other assignments, other signals, other chances. The damage is diffuse and recoverable.
A false positive, flagging an honest student, is a different category of harm. It can mean a zero on the assignment, a summons to an academic integrity hearing, a note in a permanent record, weeks of anxiety, and a rupture of trust between a student and a school that’s supposed to be on their side. For a student on a visa or a scholarship, the stakes climb higher still. And the accused often can’t prove a negative; “I wrote this myself” is genuinely hard to demonstrate after the fact. The cost of a wrong flag lands on a student in ways a missed catch never lands on the institution.
When the harm of erring one way dwarfs the harm of erring the other, the decision rule is old and clear: optimize against the worse harm. In a courtroom that principle produced “beyond a reasonable doubt.” In AI detection it means favoring precision. Better to let some AI use slide than to brand an innocent student a cheat.
A mini-scenario: the dean’s dashboard
Picture a dean handed a detector with a knob labeled, honestly for once, “sensitivity.”
Turn it to high. The end-of-term report shows 40 flagged essays across the department. The dean feels thorough. But when advisors actually read the flagged work, half of it is careful, conventional writing by strong students and non-native English speakers. The department now has 20 accusations it can’t stand behind, 20 furious families, and a reputation problem. High recall, ugly precision.
Turn it to low. The report shows 8 flags, and 7 of them, on review, are clearly AI, sloppy and generic. One is a maybe. The department has a short list it can actually investigate like adults, one conversation at a time. A few light AI users weren’t caught. High precision, imperfect recall, and a policy the dean can defend at a hearing.
The second dean sleeps better, and deserves to. The tool didn’t change. The threshold did.
The base rate makes it worse than you’d think
Precision has a nasty dependence most people miss: it depends on how many people actually cheated, not just on how good the detector is.
Suppose a detector has a 5% false-positive rate, which sounds tolerable. Now suppose only 5 students in a class of 100 used AI. The detector flags most of the 5 real cases, fine. But it also flags 5% of the 95 innocent students, which is about 5 people. Your flagged pile is now half innocent. Precision has collapsed to roughly 50%, not because the tool got worse, but because the pool of innocent students it can misfire on is so much larger than the pool of guilty ones. That’s the base-rate problem that makes a tiny false-positive rate dangerous, and it means “95% accurate” can still mean “half the accusations are wrong” in a class where cheating is rare.
The lower the true rate of AI use, the more a school should distrust its flags, and the harder it should lean on precision over recall.
What optimizing for precision looks like in practice
It’s less about a setting and more about a posture.
- Treat a flag as a question, not a verdict. A high score opens a conversation, it doesn’t close a case. Ask the student to walk you through their drafting.
- Raise the bar for action. Don’t act on a single tool’s score. Look for corroboration: draft history, sudden style shifts, an inability to discuss the work.
- Publish the policy. Tell students the detector is one signal among several and can be wrong. That honesty is itself a hedge against false-positive fallout.
- Remember who bears the error. The whole reason to favor precision is that the person who pays for a false positive isn’t the one who chose the threshold.
None of this means detection is worthless. It means the number is a probability about statistical patterns, and the humane way to use a probability is to be slow and careful in the direction where being wrong hurts most. Even OpenAI’s own classifier caught only about a quarter of AI text while flagging some human writing, and the company retired it in 2023 for low accuracy. If perfect recall was never on the table, the honest move is to protect precision and use the tool with humility.
Frequently asked questions
What’s the difference between precision and recall in AI detection?
Precision asks how many of the flagged papers were actually AI; recall asks how many of the actually-AI papers got caught. A tool can be strong at one and weak at the other. High precision means few false accusations but some AI slips by; high recall means it catches almost everything but snares innocents too.
Why can’t a detector just have both high precision and high recall?
Because a single threshold controls both. Flag more aggressively and recall rises while precision falls; flag conservatively and the reverse. Human and AI writing overlap in the middle, so no cutoff separates them cleanly, and every vendor picks a compromise point on that curve.
Which should a school optimize for?
Precision, almost always. A false negative is a missed catch; a false positive is an honest student accused of cheating, with grades, hearings, and records on the line. Those costs aren’t symmetric, so you optimize against the worse harm, which means favoring precision even when some AI use goes uncaught.
Doesn’t optimizing for precision let cheaters get away with it?
Some, yes. The alternative is accepting false accusations of honest students to catch more cheaters. Most schools, seeing the numbers, decide punishing an innocent is the worse failure. Detection can’t catch everything regardless, so precision plus human judgment is the more defensible stance.
How does the base rate change this?
A lot. When few students actually used AI, even a low false-positive rate can produce more false accusations than true catches, because the false-positive rate applies to the large innocent majority. That’s the base-rate problem, and it’s why precision, not raw accuracy, is the number that matters.
The bottom line
Precision and recall are just two names for the two ways a detector can be wrong, and a single threshold decides which mistake it makes more often. Schools should choose precision, on purpose and out loud, because the student who gets falsely accused pays a price no missed catch ever imposes on the institution. Use the score to start a conversation, never to end one. If you want to see how a detector reads a specific piece before it becomes someone’s problem, run a draft through the free checker, browse more on how detection works, or see the plans.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.


