The Accuracy-Over-Time Story Nobody Tells: Detectors Get Worse on Old Text Too
05 Aug 2026
AI detection accuracy on older text decays quietly in both directions: retrained detectors gradually forget the statistical tells of yesterday’s models, and they increasingly misread older human writing — formal, convention-bound, pre-ChatGPT prose — as machine-made. The industry’s entire accuracy conversation points forward: can detectors keep up with the newest models? Nobody benchmarks the other direction, which is how a strange fact goes unmeasured — the same essay, unchanged since 2021, can earn different verdicts from different detector versions, and a document’s odds of being flagged can drift for years after the last word was written. For anyone scanning archives, rechecking old submissions, or defending work they wrote before the AI era, the untold half of the accuracy story is the half that matters.
Key takeaways
- Detectors are moving instruments: every retraining shifts the decision boundary, so verdicts on unchanged text change over time.
- Old AI output fades from training mixes tuned against current models — and cross-generator benchmarks show detectors falter on generators they weren’t tuned for.
- Old human writing gets the opposite problem: formal, regular, convention-bound prose scores machine-like, which is how the U.S. Constitution got flagged.
- Pre-ChatGPT text flagged by modern detectors is the cleanest false-positive proof possible — the accused model didn’t exist yet.
- No vendor publishes accuracy against dated corpora, so retroactive scanning runs at unknown error rates.
The forgotten direction of drift
Every published accuracy number is a snapshot: this detector version, against these generators, on this day. The industry treats staleness as a one-sided problem — models improve, detectors chase — and the chase is real. But a classifier doesn’t just learn new boundaries; it *moves* its old ones. Each retraining redraws the line between “human-like” and “machine-like” to fit the current adversary, and every document on the wrong side of a moved line changes verdicts without changing a word.
This is ordinary machine-learning drift, but detection has a property that makes it consequential: the inputs don’t expire. Essays sit in institutional archives. Articles stay published. Turnitin retains submissions for years. Text written once gets *re-evaluated* under regimes that didn’t exist when it was written — and no one measures how the instruments perform on the past, because every benchmark, vendor or academic, tests against current generators. Accuracy on older text is an unmonitored quantity in a system that keeps making judgments about older text.
AI detection accuracy on older text: why it decays
Old machine text falls out of the training mix. A detector tuned against today’s frontier output learns today’s tells. The stiff cadences of GPT-3-era prose — obvious in 2021, quaint now — contribute less and less to the training signal as corpora refresh, and recognition of them atrophies. The RAID benchmark quantified the general mechanism: detector performance drops sharply on output from generators outside the tuning set, before any adversarial editing. “Outside the tuning set” is precisely what last generation’s models become. We’ve traced how much the target has moved in the generational gap in AI writing from GPT-2 to today — a detector chasing the newest voice is, by construction, walking away from the older ones.
Thresholds get re-tuned for a different world. As frontier models close the statistical gap with human prose, vendors adjust decision thresholds to hold headline accuracy — and a threshold tuned for subtle modern output behaves differently on the coarse output and clean human baselines of an earlier era. The operating point that’s optimal against 2026’s adversary was never validated against 2021’s documents.
Older human writing looks like the enemy. The false-positive side is worse. Statistical detectors flag predictability: even rhythm, conventional vocabulary, low surprise. That profile describes machine output — and also describes formal writing from any era before “write like a human” meant “write casually.” Historical documents, legal boilerplate, catechism-style academic prose: all of it scores eerily regular. The canonical embarrassment is well documented — detectors flagged the U.S. Constitution as AI-written, a result Ars Technica used in 2023 to illustrate why perplexity-based verdicts fail on famous, formal, much-imitated prose. There’s a cruel loop underneath: the models *trained on* centuries of human writing, so the older canonical styles are literally part of what “AI-like” now sounds like. A detector hunting model echoes is partly hunting the human prose the models learned from.
The reference class rots. Detectors also need a “human” baseline, and as we covered in how training-data contamination corrodes detection accuracy, post-2023 human corpora are salted with AI-assisted prose. As the baseline drifts toward the machine distribution, genuinely old human text — cleaner than the contaminated baseline, but differently regular — sits further from *both* classes the detector knows, and its verdicts get less predictable.
Where this bites: the retroactive scan
The decay would be trivia if old text stayed unjudged. It doesn’t. Institutions rescan archived submissions when new tooling arrives; disputes reopen old documents under new detector versions; hiring and admissions officers run historical writing samples through current tools. Each scan silently assumes the instrument is calibrated for the document’s era. None is.
The sharpest case is pre-ChatGPT text flagged as AI — an essay from 2019 accused of being written by a model that didn’t exist. These cases are gold for one purpose: they’re false positives with *proof*, no statistical argument required, and they should permanently calibrate how much confidence any score deserves. But the same drift cuts subtler and meaner day-to-day: a student’s 2022 work rescanned in 2026 under a boundary tuned for 2026 models; an archive audit that flags the most formally written — often the most non-native — old submissions, compounding the bias Liang et al. documented. A verdict that changes when the text didn’t is telling you about the detector. Institutions consistently hear it as telling them about the writer.
What this means in practice
If you steward archives or old submissions: don’t rescan the past with present instruments and act on the output — no vendor publishes error rates against dated corpora, so you’d be enforcing at unknown accuracy. If a scan happens anyway, require date evidence before any accusation; a timestamp that predates the accused technology ends the conversation. If you’re a writer defending old work, that’s your lever — version history, submission records, drafts — and it’s absolute in a way no statistical argument is. And for text you’re producing *now*, remember that today’s clean scan is one detector-version’s opinion: what protects you over time isn’t a score, it’s process evidence plus knowing how your prose reads at the sentence level — try it on your own text to see that view, or see what each plan handles if you check documents at volume. For more on how these instruments age, browse the rest of our writing on detection.
Frequently asked questions
Do AI detectors still recognize text from older models like GPT-3? Less reliably than you’d assume. Detectors are retrained to chase current models, and their decision boundaries shift with each update — the tells of 2021-era output fade from training mixes dominated by newer generations. Cross-generator evaluations like RAID show detectors degrading sharply on model output they weren’t tuned for, and yesterday’s models quietly join that unfamiliar category. Recognition of old AI text is nobody’s benchmark, so nobody notices it slipping.
Why do detectors flag old human writing as AI? Because formal, polished, convention-bound prose has exactly the statistical profile detectors call machine-like: predictable vocabulary, even rhythm, low surprise. Historical documents, legal boilerplate, and traditional academic writing all score this way — famously, detectors flagged the U.S. Constitution as AI-written. Older prose often follows stricter conventions than modern informal writing, so the further back a human text comes from, the more its regularity can resemble a language model’s.
Can an essay written before ChatGPT be flagged by today’s detectors? Yes, and such cases are the cleanest proof of false positives in existence — a 2021 essay physically could not have been written by a 2023 model. Pre-AI text carries no immunity: detectors read statistical texture, not timestamps, and formal student writing from any era can score machine-like. If old work is ever rescanned under a newer detector version, the verdict can differ from the original scan of the very same words.
What is training-data contamination and how does it affect old text? It cuts both directions across time. New detectors train on human corpora increasingly salted with AI-assisted prose, which blurs the human reference class they compare everything against — including old text. Meanwhile the models themselves trained on decades of human writing, so classic prose styles literally shaped what AI output sounds like; a detector hunting AI patterns is partly hunting echoes of the older human writing the models learned from.
What does detector drift mean for institutions scanning archives? That retroactive scanning is close to indefensible. A score produced today reflects today’s decision boundary, tuned against today’s models — applied to text written under a different statistical regime, its error rates are unknown and unmeasured, since no vendor benchmarks against dated corpora. Any institution rescanning archived work should require date-stamped process evidence before acting on a flag, and treat a changed verdict on unchanged text as evidence about the detector, not the writer.
The bottom line
Detection accuracy is always narrated as a race against the newest models, and the narration hides the cost of running: every stride away from yesterday’s adversary is also a stride away from yesterday’s documents. Old machine text slips out of recognition; old human text — formal, regular, echoed in the models’ own training — slips *into* suspicion; and none of it shows up in benchmarks, because nobody tests against the past. The practical rule falls out cleanly. A detector’s verdict is dated the moment it’s issued — trustworthy, to whatever degree it ever is, only for the era it was tuned in. Text is permanent. Instruments aren’t. Any policy that forgets which one changes will keep discovering AI in documents written before AI could write.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.
