AI Detection for HR and Recruiters: Spotting AI-Written Cover Letters and Resumes Fairly
13 May 2026
A recruiter opens forty cover letters before lunch. Three of them have a certain sheen to them, the kind of tidy, balanced prose that makes you wonder if a chatbot wrote it. So you paste one into an AI detector, it spits back “92% AI,” and now you’re holding a number you don’t quite know what to do with.
That number feels like evidence. It isn’t. This piece walks through what AI detection can and can’t tell you about job applications, where it goes badly wrong, and how to use it without quietly building an unfair hiring process.
Key takeaways
- AI detectors output a probability, not proof. A high score on a cover letter is a reason to read more carefully, not a reason to auto-reject.
- Resumes are short, formulaic, and keyword-heavy by design. That’s exactly the kind of text detectors misread, so resume scores are close to meaningless.
- Detectors have been shown to flag non-native English writers more often. Using them as a filter can create a discrimination problem you won’t see coming.
- Most candidates already use AI to polish applications, and that’s usually fine. The real question is whether the person can do the job.
- If you screen at all, do it openly: state your policy, treat detection as one input, and keep a human deciding every case.
What an AI detector actually measures
AI detectors don’t read for meaning. They measure statistical fingerprints. Two of the big ones are perplexity (how predictable each word is, given the words around it) and burstiness (how much sentence length and rhythm vary). Machine-generated text tends to be smooth and predictable; human writing tends to wobble. Detectors look for that smoothness and convert it into a percentage.
That percentage is a guess about probability, not a fact about authorship. It’s worth understanding the mechanics if you’re going to lean on these tools at all, because the same math that makes them sometimes useful also makes them fail in predictable ways. We dug into the perplexity and burstiness statistics separately if you want the deeper version.
The honest baseline: even OpenAI pulled its own AI text classifier in July 2023 because of low accuracy. When OpenAI tested it, the tool correctly flagged only about a quarter of AI-written text while wrongly tagging some human writing as machine-made. The company that builds the model that worries you couldn’t reliably detect its own output. That should set your expectations for the third-party tools claiming they can.
Why cover letters are a hard target
Cover letters are short. Most run 200 to 400 words, and detectors get shakier the less text they have to work with. A four-paragraph letter doesn’t give the statistics much room to stabilize, so scores bounce around.
They’re also a genre that rewards a particular polish. Candidates have been taught for decades to write cover letters in a measured, professional register. That register, careful, balanced, a little formal, happens to look a lot like what a language model produces. A genuinely human applicant who’s a strong writer can score “AI” simply for writing well.
Then there’s the obvious wrinkle: most applicants now use AI somewhere in the process. They draft with it, edit by hand, run it past a friend, and send. A lightly assisted human letter and a lightly edited AI letter can land in the exact same score range. The detector can’t separate “wrote it with help” from “didn’t write it at all,” and for hiring purposes those are very different things.
Resumes: basically a detector’s worst case
If cover letters are hard, resumes are nearly hopeless. They’re built from sentence fragments, standard section headers, and recycled industry phrasing (“cross-functional collaboration,” “drove a 30% increase”). There’s no flowing prose for perplexity to chew on, and bullet points kill burstiness because they’re structurally uniform.
The result is that resume AI scores swing wildly and mean almost nothing. A hand-typed resume from a fifteen-year veteran can read as “AI” because the language of resumes is, by design, predictable and templated. Treat any AI percentage on a resume as noise.
The fairness problem you can’t ignore
Here’s where this stops being a quality issue and becomes a risk issue. A 2023 study led by Weixin Liang, James Zou, and colleagues at Stanford found that several popular detectors flagged writing by non-native English speakers as AI-generated far more often than writing by native speakers. In their tests on TOEFL essays written by non-native writers, the detectors misclassified more than half the samples on average, while handling native-speaker writing accurately. The likely reason: non-native writers often use simpler vocabulary and more predictable constructions, which detectors read as “machine-like.”
Sit with what that means in a hiring context. If you filter applications by AI score, you may be systematically downranking candidates based on their first language. That’s the shape of a discrimination claim, and “the software flagged them” is not a defense anyone wants to give in front of a regulator or a judge. Detection bias quietly converts into hiring bias.
A quick scenario
Two candidates apply for a marketing coordinator role. The numbers below are illustrative, but the pattern is the one the research describes.
Priya, who learned English as her third language, writes a clear, plain cover letter. Her sentences are short and even. The detector returns a high “AI” score.
Mark, a native speaker, runs his real experiences through ChatGPT and asks it to “make this sound natural,” then tweaks a few lines. The detector returns a low one.
If your process auto-rejects above a fixed threshold, you just cut the human writer and advanced the AI-assisted one, while introducing a language bias. The score told you nothing useful about either person’s ability to coordinate marketing campaigns.
So what should HR and recruiters actually do?
You don’t have to ban detection. You have to stop treating it as a verdict. A few principles keep it sane and fair.
Use detection as a flag, never a filter
A high score can earn an application a closer human read. That’s the ceiling of its usefulness. It should never trigger an automatic rejection, and it should never be the deciding factor on its own. One weak signal does not outweigh experience, interviews, and references.
Be transparent about your policy
If you care about AI use, say so in the job posting. “We welcome AI-assisted applications, but the interview includes a short writing exercise” is honest and sets expectations. Secret detection that quietly tanks applications is the worst of both worlds: it’s unfair to candidates and indefensible if challenged.
Test the skill you actually care about
If unassisted writing genuinely matters for the role, measure it directly. A short, timed, supervised writing exercise during the interview tells you more than any detector reading of a take-home document. You’ll see how the person thinks under real conditions instead of guessing from a probability score.
Keep a human in every loop
Document your process so decisions are consistent. Detection output goes to a person, the person makes the call, and the reasoning gets recorded. That’s both fairer and far more defensible than an algorithm silently sorting your pipeline.
Reframing the real question
Step back and ask what you’re actually worried about. It’s usually one of two things: that the candidate misrepresented their abilities, or that they can’t write well enough for the job. Neither is something a detector answers.
Misrepresentation is caught by reference checks, work samples, and probing interview questions, not by sniffing for ChatGPT. Writing ability is caught by watching them write. If a candidate used AI to organize honest accomplishments into a clean letter, that’s roughly the same skill as knowing when to ask a colleague to proofread. Plenty of excellent employees work exactly that way.
The candidate who pastes a generic, fabricated letter usually gives themselves away the moment you ask a specific question about their experience. A detector isn’t doing the work there; your interview is.
The bottom line
AI detection can be a small, honest input into hiring. It cannot be the gatekeeper. The scores are probabilities, they’re worst at exactly the documents you’re screening, and they carry a documented bias against non-native speakers that can turn into a fairness problem fast. Use them lightly, tell candidates what you’re doing, and judge people on whether they can actually do the job.
Want to see how unstable these verdicts get before you build a hiring policy around them? Paste in a few real human letters, a few AI ones, and some non-native English writing, and watch the scores jump. You can run a real cover letter through a checker and see the per-sentence breakdown for yourself, which does more to calibrate your skepticism than any vendor claim. There’s also more on AI writing and detection if you want to go deeper before setting a standard for your team.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.
