How to Document an AI Suspicion Properly Before You Report a Student
12 Jul 2025
A detector flags an essay at 98% AI. Your stomach drops, you reach for the report form, and you stop. Good. The moment between suspicion and accusation is where careful documentation either saves you or sinks you, and most of us were never taught how to handle it.
This guide walks through documenting an AI cheating suspicion the right way: what to collect, what to leave out, and how to keep a record that holds up if it ever reaches an integrity board. The goal isn’t to “win” a case. It’s to make sure that if you do report a student, you’re standing on evidence instead of a number.
Key takeaways
- A detector score is a starting signal, not evidence. Write down what made you suspicious before you act, and keep that separate from your gut reaction.
- Build a file: the prompt, the submission, the detector output (tool, date, version, exact score), and specific notes on voice, sources, or process.
- Detectors return probabilities, and they misfire more on non-native English writers. Note those limits in your own record so you don’t overstate what you have.
- Process evidence you already collect, like draft history and in-class samples, tells a stronger story than any single percentage.
- Talk to the student before filing. The conversation often answers the question, and it belongs in your documentation regardless.
Why documentation matters more than the score
Here’s the uncomfortable truth: the detector did the easy part. It gave you a number. Everything that turns that number into a fair, defensible decision is on you.
AI detectors don’t catch cheating. They estimate how statistically similar a piece of text is to patterns common in machine-generated writing. That’s why detectors return a probability, not proof — a 95% reading means “this looks like the kind of thing models produce,” not “this student used a model.” Those are different claims, and the gap between them is exactly where students get hurt.
We also know the tools are imperfect in specific, documented ways. A 2023 Stanford study led by Weixin Liang, published in the journal *Patterns*, found that detectors misclassified more than half of TOEFL essays written by non-native English speakers as AI-generated, while judging essays by native US students almost perfectly. OpenAI pulled its own AI Text Classifier off the market in July 2023, citing a low rate of accuracy. If the company that builds the models couldn’t reliably detect them, a single third-party score deserves humility.
So documentation does two jobs. It records what you genuinely observed, and it protects a student from a decision made on a bad afternoon and a screenshot.
What to collect, step by step
Think of yourself as building a small case file, not winning an argument. Keep it factual and dated. Here’s the order that works.
1. Capture the assignment context first
Before you look at anything else, save the assignment prompt, the due date, the submission timestamp, and any instructions you gave about AI use. This sounds boring. It’s the part people skip and regret. If your syllabus said “you may use AI to brainstorm but not to draft,” that line shapes everything that follows.
2. Save the submission exactly as received
Download or print the actual file the student turned in. Don’t paraphrase it later from memory. If it was submitted through an LMS, note the platform and keep the original. You want the unedited artifact, because details you ignore today — a citation format, an odd phrase, a metadata timestamp — can matter next week.
3. Record the detector output in full
Screenshot the result, but write down the parts a screenshot misses: which tool you used, the version or date if shown, the exact score, and the date and time you ran the check. If the tool highlights specific sentences, save that view too. One number with no context is weak. “GPTZero, run June 14, flagged paragraphs 2 and 4 at high probability” is something you can actually discuss.
If you only have a number from one tool, run the same submission through a second detector and compare. Detectors disagree with each other constantly, and a reading that contradicts the first is itself useful evidence — in the student’s favor or yours. If you’re weighing which checker to lean on, it helps to know how each one prices access and what it actually reports before you trust a single score.
4. Write down your specific observations
This is the heart of good documentation, and it’s where most files fall apart. Don’t write “this feels AI.” Write what you noticed:
- “The voice is flat and uniform, unlike the student’s three earlier discussion posts, which used contractions and personal asides.”
- “Cites a 2024 study that doesn’t appear to exist; the DOI returns nothing.”
- “Sentence length barely varies for four paragraphs, then changes abruptly in the conclusion.”
Concrete, checkable observations survive scrutiny. Impressions don’t.
5. Pull whatever process evidence you have
Draft history, Google Docs version timelines, in-class writing samples, outline submissions, peer-review comments — anything that shows how the work came to be. A document that materializes fully formed at 11:58 p.m. with no edit history tells a different story than one with 40 saved revisions. Process evidence is often more persuasive than any detector, because it’s about this student’s actual behavior, not a statistical average.
A short scenario: doing it right
Mr. Alvarez teaches sophomore English. A student’s essay scores 91% AI on the school’s detector. His first instinct is to email the integrity office. Instead, he opens a document and starts a file.
He saves the prompt, the submission, and the detector screenshot with the date. Then he pulls the student’s two earlier graded essays and notices the new one reads nothing like them — no rhetorical questions, no slang creeping into the analysis, none of the student’s usual habit of overusing semicolons. He checks the three sources cited; two are real, one isn’t. He notes all of this in plain language.
He also writes a line he’s tempted to leave out: “Student is a recent arrival and a non-native English speaker; detector bias is a known issue per the Stanford study. Score alone is not conclusive.” That sentence makes his file more credible, not less, because it shows he weighed the obvious counter-argument.
Then he emails the student to set up a quick chat about the essay’s sources. That conversation — not the 91% — is what will actually settle it. Whatever happens, Mr. Alvarez has a record built on observations, context, and fairness.
What to leave out of your documentation
Just as important as what you include:
- Conclusions stated as fact. Write “I observed X,” not “the student cheated.” Your file records evidence; the integrity process reaches the verdict.
- Emotional language. “Blatant,” “obvious,” “lazy” — cut all of it. Tone leaks into the record and undermines you.
- A single tool treated as a verdict. If your whole case is one percentage, you don’t have a case yet.
- Comparisons to other students. Each suspicion stands on its own facts.
- Anything you didn’t actually verify. If you suspect a fabricated source, check it before you write it down.
Loop in the student before you escalate
Unless your institution’s policy dictates otherwise, a direct, low-stakes conversation should usually come before a formal report. Ask the student to walk you through their process: where they started, what sources they used, how they’d explain a particular claim. Honest writers can almost always reconstruct their own work. The conversation frequently resolves the whole thing, and when it doesn’t, your notes from it become part of the file.
Keep that talk factual and non-accusatory. You’re gathering information, not delivering a sentence. For the broader picture of fair detection practices, there’s more on the educator side of AI detection worth reading before you build your own classroom approach.
Frequently asked questions
Is an AI detector score enough to report a student? No. A detector estimates the probability that text resembles AI writing — it doesn’t prove a specific student used AI. Treat the score as a prompt to look closer, then build your case from the assignment, the student’s writing history, and a conversation. A screenshot of a percentage is not, by itself, a report.
What should an AI suspicion file actually contain? The prompt and submission date, the full submission, the detector output (tool, version, date run, exact score), and your written observations about what stood out. Add process evidence like draft history where you have it. Record what you observed, not what you concluded.
Can a high detector score be wrong? Yes, and more often than people expect. The 2023 Stanford study found detectors misclassified more than half of non-native English writers’ TOEFL essays, and OpenAI retired its own classifier in July 2023 for low accuracy. Formulaic structure and heavy grammar-tool use can push an honest essay’s score up.
Should I tell the student I’m suspicious before reporting? Usually yes, unless policy says otherwise. A conversation about their process often clears things up, and if it doesn’t, it strengthens your file. The student deserves a chance to respond before a formal accusation hits their record.
How long should I keep this documentation? Follow your institution’s retention policy, which typically covers at least the appeal window. Keep dated, factual notes even when a suspicion resolves informally — patterns across assignments are easier to see when you have a record.
Before you hit “report”
Slow down enough to build a file that’s about evidence, not anxiety. A careful record protects honest students from bad numbers and gives genuine cases the weight they need to be taken seriously.
If you want to understand what a detector is really telling you before it ever lands on a student’s desk, run a draft through PaperBleach and read the result for what it is — a probability, not a verdict.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.
