logo

How the Definition of ‘AI-Generated’ Keeps Shifting and Breaking Detection

Paperbleach
Paperbleach

05 Aug 2026

Nobody can say what counts as AI generated text anymore — not regulators, not universities, not the detector vendors — because the category has quietly fractured into a spectrum running from autocomplete to full ghostwriting, with every institution drawing its line in a different place. That fracture isn’t a pedantic complaint. Detection tools are classifiers, and classifiers need a stable target: when the definition of the thing being detected shifts underneath them, training data gets mislabeled, benchmarks stop agreeing with each other, and a binary verdict gets stretched over documents that are genuinely neither. The definition problem, more than any single technical failure, is why AI detection keeps breaking — and it’s getting worse.

Key takeaways

  • “AI-generated” began as a clear category in 2022 — text a chatbot wrote — and has since dissolved into a spectrum of assistance levels.
  • Regulators, schools, journals, and detector vendors all define the category differently, so one document can be AI-generated and not, simultaneously.
  • AI moved inside the writing tools themselves — keyboards, Docs, Word, Grammarly — making “no AI touched this” nearly unclaimable for ordinary prose.
  • Detection breaks because classifiers need stable labels: hybrid documents poison training data and make benchmark accuracy incomparable.
  • The practical defense is process evidence and disclosure in your policy’s own terms, plus knowing how your prose reads statistically.

What counts as AI generated text? Ask four institutions, get four answers

In December 2022 the category seemed self-evident: you prompted ChatGPT, it wrote an essay, that essay was AI-generated. Every definition since has been an attempt to hold that clarity while the ground moved.

The EU AI Act’s Article 50 defines the category operationally — content produced by a generative system should carry machine-readable marking — which locates “AI-generated” at the moment of *production*, whatever happens afterward. Turnitin’s documentation defines it statistically: their detector flags prose whose patterns resemble language-model output, which locates the category in the *finished text*, regardless of process. *Nature*’s ground rules define it contractually — LLMs can’t be authors and their use must be documented — locating it in *disclosure*. And a typical university policy defines it pedagogically, distinguishing brainstorming (usually fine) from drafting (usually not), locating it in *which cognitive work was outsourced*.

Four institutions, four locations for the same boundary: production, pattern, disclosure, process. A student who drafts their own essay and has a model tighten every paragraph produces text that is AI-generated under Turnitin’s definition, compliant under Nature’s, unmarked under the AI Act’s, and contested under the university’s. None of these authorities is being careless. They’re regulating different harms, so they’ve drawn different lines — and the detector is the only one of the four that gets asked for a yes-or-no answer.

Three forces that keep moving the line

AI moved into the tools. The 2022 picture assumed a clean separation: your word processor over here, the chatbot over there, a deliberate copy-paste between them. That separation is gone. Autocomplete finishes sentences in Gmail; Word and Docs draft and rewrite natively; Grammarly rebuilds whole sentences with the same transformer technology detectors are trained to catch. When generation is a suggestion popup inside every text field, “no AI touched this document” becomes something almost no ordinary writer can claim — and the interesting question shifts from *whether* to *how much and which parts*, which is precisely the question a binary label can’t hold. We drew this boundary in more detail in the difference between editing with AI and generating with AI.

Usage normalized faster than policy. Stanford-led corpus work found measurable LLM fingerprints in scientific abstracts within a year of ChatGPT’s launch — most of it language assistance by researchers writing in their second language, and much of it now explicitly permitted with disclosure. Each time an institution legalizes a use case, text that was “AI-generated” in the accusatory sense gets redefined as ordinary assisted writing, and yesterday’s contraband becomes today’s baseline. The category keeps losing territory to its own normalization.

Iteration erased the author boundary. Real workflows are now conversational: human outline, model draft, human restructure, model polish, human final pass. Ask which sentences of the result are “AI-generated” and the honest answer is a shrug — provenance interleaves at the clause level. Even Grammarly, whose business is writing assistance, effectively conceded the finished text can’t answer the question: its Authorship feature records *how* a document was composed — typed, pasted, generated — because inspecting the artifact afterward no longer settles anything.

Why detection breaks when the category wobbles

A supervised classifier is only as coherent as its labels, and every mechanism above corrupts the labels.

Training corpora rot first. “Human” reference text scraped after 2023 is saturated with AI-polished prose, and “AI” samples increasingly mean human-edited hybrids — so the two distributions detectors learn to separate now overlap by construction. Benchmarks inherit the incoherence: one evaluation files lightly edited AI text as machine-written, another as human, and the same detector scores brilliantly on the first and terribly on the second without changing a line of code. Accuracy claims stop being comparable because they encode different answers to the definition question.

Then the verdict itself stops meaning anything. When a detector says “82% AI,” is that a probability of full generation, an estimate of the machine-touched fraction, or resemblance to model style? Against a spectrum of real documents, both flagging and clearing a hybrid is defensible — which makes every score unfalsifiable and every dispute unresolvable. We’ve catalogued how tools handle the gray zone in whether heavily edited AI text still counts as AI; the honest summary is that each vendor answers a differently worded question and reports it on the same-looking scale.

What this means in practice

For writers, the ambiguity cuts both ways: it means good-faith assistance can be flagged, and it means the burden of clarity falls on you. Read the operative policy and note *which* definition it uses — production, disclosure, or process. Disclose in that policy’s own vocabulary. Keep drafts and revision history, because process evidence answers the question detectors can’t. And know your statistical shadow: whatever your policy says, the pattern-based definition runs anyway, so try it on your own text and see the sentence-level read before a gatekeeper does — or see what each plan handles if you’re checking work at volume.

For institutions, the lesson is to stop asking a pattern instrument to enforce a process category. Define the violation in terms of undisclosed outsourcing, collect process evidence, and treat a detector score as a reason to ask questions, never as an answer. For the longer map of this territory, browse the rest of our writing on detection.

Frequently asked questions

What officially counts as AI-generated text? There is no official definition — only overlapping ones that disagree. The EU AI Act treats content produced by a generative system as needing machine-readable marking. Turnitin says its detector identifies prose whose patterns match language-model output, however it was used. Journals define the category by disclosure obligations, and universities each draw their own line between assistance and generation. The same document can be AI-generated under one definition and not under another.

Does using Grammarly or autocomplete make my text AI-generated? Under most policies, no — but under some detectors, sometimes yes. Suggestion-level tools like autocomplete and grammar fixes are broadly treated as assistance, not generation. The complication is that modern versions rewrite whole sentences with the same underlying model technology detectors are trained to spot, so heavily accepted suggestions can nudge prose toward machine-typical patterns. Policy answers and statistical answers to this question have quietly diverged.

Why do institutions define AI-generated differently from detectors? Because they’re answering different questions. A policy defines a violation, which requires intent and process: who conceived the ideas, who drafted, what was disclosed. A detector measures a statistical resemblance in the finished prose and knows nothing about process. So institutions write process-based rules while running pattern-based tools, and the mismatch — process categories enforced by a pattern instrument — is where most disputes start.

How does the shifting definition break AI detection? Classifiers need a stable target, and this one moves on both ends. Training data gets mislabeled — human corpora now contain AI-polished prose, and hybrid documents get filed as one or the other. Benchmarks disagree about which category edited text belongs in, so accuracy numbers stop being comparable. And a binary verdict compresses a spectrum: when most real documents are partly machine-touched, both flagging them and clearing them is defensible, which makes any score unfalsifiable in practice.

How should writers protect themselves given the ambiguity? Anchor to the strictest definition in your context, then keep evidence. Read the actual policy language of your institution or client — some regulate generation only, others any undisclosed assistance. Disclose in the terms that policy uses, keep drafts and revision history to document your process, and check how your finished prose reads to a detector before a gatekeeper does, since the statistical definition operates regardless of what your policy says.

The bottom line

“AI-generated” was a stable category for about a year — the year when generation lived in one chatbot window and writing lived everywhere else. Since then the boundary has been dissolved from three directions at once: the tools absorbed the models, the institutions legalized more of the spectrum, and the workflows interleaved human and machine at the sentence level. Detectors inherit the wreckage, because you cannot reliably classify a category whose definition depends on who’s asking. The shift that matters is already visible at the edges — from guessing what a text *is* to recording how it was *made*, from verdicts to disclosure. Until that transition completes, the only stable ground belongs to writers who keep their process visible and know how their prose reads to the instruments, whatever this year’s definition happens to be.

Try it on your own text

Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.