PaperBleachPaperBleach
logo

Why Detector Scores Drift Over Time as Models and Training Data Change

P
Paperbleach

06 Jul 2026

Run a paragraph through an AI detector today, jot down the score, then run the exact same paragraph through the same tool six months from now. Don’t be shocked if the number moved. Not because your writing changed — it didn’t — but because the detector did. And so did the models it’s chasing, and arguably the way people write, too.

AI detection isn’t a fixed measurement like weighing a bag of flour. It’s a moving instrument aimed at a moving target across shifting ground. That instability has a name in machine learning, and it has real consequences for anyone tempted to treat a score as durable proof.

Key takeaways

  • Detector scores aren’t reproducible over time. The same text can read human one month and AI the next.
  • Detectors chase a moving target: new models write differently, so what counts as “AI text” keeps changing — the core of concept drift.
  • Vendors quietly retrain tools, swap scoring models, and adjust thresholds, shifting scores with no change to your writing.
  • As models improve, the AI-versus-human gap shrinks, so detectors degrade unless constantly updated.
  • A score is a snapshot from one tool on one day, not a permanent fact — which is a real problem when it’s used as evidence.

A score is a snapshot, not a measurement

We’re used to measurements being stable. A meter stick reads the same length this year and next; that reliability is the whole point of a measurement. AI detector scores don’t work that way, and mistaking one for the other is where a lot of trouble starts.

A detector’s output depends on three things that all move: the model that supposedly wrote the text, the detector’s own trained parameters and thresholds, and the underlying distribution of how humans and machines write. Change any of them and the score for a fixed piece of text can change. Since all three drift more or less continuously, a detector score is best read as a snapshot — this tool, this version, this date — rather than a fact about the document that will hold up when you check again.

Concept drift: the target keeps moving

The deepest source of drift has a proper name: *concept drift*, the machine-learning term for when the thing you’re detecting changes shape while your detector stays still.

A detector — especially a trained classifier rather than a zero-shot method — learns what AI writing looked like *at training time*. It internalizes the tells of the models that existed then: certain phrasings, a particular flatness, characteristic transitions. Then the world moves. New model versions ship every few months, each writing a little differently, often more naturally. “AI text” in 2026 is not the same object it was in 2023.

The detector, frozen on the old patterns, is now judging new output with outdated expectations. It doesn’t announce that it’s confused; it just quietly gets things wrong more often. Its accuracy erodes not because anyone broke it, but because the concept it learned went stale. This is exactly why detectors need constant retraining just to tread water — and even a zero-shot method like DetectGPT’s curvature approach isn’t immune, since it works best against the model that wrote the text, and that model keeps being replaced by newer ones it wasn’t tuned for.

The vendor changes the tool underneath you

Concept drift is the slow, structural cause. There’s a faster, blunter one: the detector company changes the product.

Detector vendors update their tools all the time, and rarely with a changelog you can read. They retrain on fresh data, swap in a different scoring model, or nudge the threshold that separates “human” from “AI” to chase better numbers. Every one of those changes can move the score for text that never changed at all. A document that read comfortably human last spring can trip the wire this fall because the vendor tightened the threshold, not because a single word is different.

This is what makes detector results non-reproducible in the way that matters most. If you flagged an essay in March and can’t recreate that flag in September, the flag was never a stable finding — it was a reading from a version of a tool that no longer exists. For anyone using detectors in a high-stakes setting, that alone should give pause.

The gap keeps shrinking

Layer on the long-term trend and it gets harder still. Detectors work by exploiting a gap between how machines write and how humans write. As models improve, that gap narrows — the “AI cloud” of writing patterns keeps drifting toward the “human cloud,” because making output more human is precisely what model developers are optimizing for.

So even a perfectly maintained detector faces a target getting closer to human every year. Text that was easy to flag from an early model becomes genuinely ambiguous from a current one. This isn’t hypothetical hand-wringing: OpenAI launched its own AI Text Classifier in early 2023 and pulled it that July, citing a low rate of accuracy — an admission, from the people who understood the models best, that reliable detection was slipping out of reach even then. The direction of travel has not reversed since.

Humans move too

Drift isn’t only about the AI side. The human baseline shifts as well, and a detector calibrated on older human writing can misread current writing for that reason alone.

Language conventions change — new slang, new formats, evolving norms about tone and structure. And there’s a genuinely strange feedback loop now underway: as people read and absorb AI-assisted writing every day, some of its cadences seep into how humans write, blurring the line further from the other side. Meanwhile the reference data a detector was calibrated on can already be skewed — recall the finding from Liang and colleagues that detectors were biased against non-native English writers, a baseline problem that only compounds as the human distribution moves. A detector frozen on yesterday’s human writing is misjudging today’s from both directions at once: the AI target moving *and* the human baseline moving.

What to actually do about it

The practical response isn’t to find the one detector that doesn’t drift — there isn’t one. It’s to hold detector scores with the right amount of grip.

Treat any score as perishable. If a result matters, write down the tool, its version if you can find it, and the date, and accept up front that you may not be able to reproduce it later. Never let a single flag stand as a self-sufficient verdict, especially in an academic-integrity context where the stakes are someone’s record. Weigh it beside things that *don’t* drift — the writer’s drafts, their revision history, their sources, their ability to talk through their own work. Drift is one more entry on the long list of reasons a detector is a hint worth investigating, not a judgment you can bank.

Frequently asked questions

Why did the same text get a different AI score months later?

Because the detector isn’t a fixed measuring stick — it changes underneath you. Vendors quietly retrain their tools, swap the scoring model, and adjust thresholds, so the same paragraph can score human one month and AI the next with nothing about your writing changed. Detector scores aren’t reproducible over time the way a ruler is, which is a serious problem when they’re treated as stable evidence.

What is concept drift in AI detection?

Concept drift is when the thing you’re detecting changes shape while your detector stays still. A detector learns what AI text looked like at training time. Then new models ship that write differently, so “AI text” becomes a moving target. The detector, frozen on yesterday’s patterns, gradually gets worse at recognizing today’s output — its accuracy erodes because the world moved and it didn’t.

Do detectors get worse as language models improve?

Generally yes. As models get better at varied, human-like writing, the statistical gap detectors rely on shrinks — the AI cloud drifts toward the human cloud. Text that was easy to flag from a 2022 model can be much harder to flag from a 2026 one. Detectors have to keep retraining just to hold steady, and even then they’re chasing a target getting closer to human.

Does human writing changing also cause drift?

It contributes. Conventions shift — new slang, new formats, and, ironically, people absorbing habits from the AI-assisted writing they read daily. As the human baseline moves, a detector calibrated on older human text can misjudge current writing. Drift comes from both sides at once: the AI target moving and the human baseline moving, neither of which a frozen detector accounts for.

What should I do given that scores drift?

Stop treating any single score as durable proof. A flag is a snapshot from one tool on one day, not a permanent fact. If a score matters — in an academic case, say — record the tool, version, and date, understand it may not reproduce later, and weigh it alongside drafts, process, and the writer’s account rather than as a standalone verdict.

The short version

An AI detector is a moving instrument aimed at a moving target over shifting ground. The models it chases keep changing (concept drift), the vendors keep retraining and re-thresholding their tools, and the human baseline keeps moving too — so the same text can honestly score human today and AI next month. Add the long-run shrinking of the gap between machine and human writing, and no score is durable. Record it, date it, and never let it stand alone as proof. Curious where your writing sits right now? Run a draft through the free checker, then browse more detection guides or see the plans.

Try it on your own text

Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.