logo

AI Detection for Scientists Writing for the Public: Press Summaries That Don’t Read as Generated

Paperbleach
Paperbleach

30 Jul 2026

A research summary reads as human when it contains what only the researcher has — the surprise, the limitation, the three years of Tuesday-night telescope time — and reads as generated when it’s assembled from the abstract, which is why AI-drafted press summaries all sound alike and why readers have started to notice. AI detection in science communication is less about any tool’s verdict than about that reader reaction: public science writing is a trust instrument, and recognizably synthetic prose spends trust the institution took decades to earn.

This one’s personal for research institutions in a way the marketing version of this problem isn’t. Let’s take it apart.

Key takeaways

  • Press summaries flag on detectors because they blend two formulaic registers — institutional boilerplate and explainer prose — and because, increasingly, they really are AI-drafted from abstracts.
  • The credibility mechanism: readers who notice generated language infer automated scrutiny of the claims themselves. Unfair, but operative.
  • AI-assisted drafting of lay summaries is broadly permitted (journal authorship rules like Nature’s govern the papers, not the outreach) — but the scientist must verify every simplified claim, because simplification is where overclaiming hides.
  • The human elements no model can supply from an abstract: surprise, honest limitation, process detail, first-person puzzlement, and a quote the scientist actually said.
  • A detector pass is a useful register check on releases — not to chase a low score, but to catch the paragraph that must carry human signal and doesn’t: the quote.

Why the press-summary register was always going to flag

Run a decade of pre-ChatGPT university press releases through a detector and plenty would score AI-likely. The genre was formulaic before models existed: “Researchers at X University have discovered…”, “The findings, published in Y, could pave the way for…”, “‘We are excited about the potential implications,’ said Dr. Z.” Institutional review sanded every release toward the same safe shape, and safe shapes are statistically predictable — which is the only thing detectors measure. We’ve covered this false-positive mechanic for the PR industry in AI detectors and press releases.

What changed is that the register became *earned* suspicion. Drafting a lay summary from an abstract is genuinely one of the tasks language models do well, so press offices adopted it at scale — and now the formulaic register and the generated register have merged into one sound. Journalists who receive forty releases a day recognize it instantly. So do growing numbers of readers. The detector’s verdict and the reader’s verdict have converged: *this could have come from anywhere.*

For most content, that’s a rankings problem. For science communication, it’s worse.

The trust mechanism, spelled out

Public-facing science writing has one product: justified confidence that the institution behind the words did careful work. Survey research on trust in science — Pew has tracked this for years — shows it’s meaningfully lower than scientists would like and sensitive to perceived corner-cutting.

Now watch what a generated-sounding summary does to a lay reader. They can’t evaluate the methods. Their only handle on the work’s care is the care visible in how it’s communicated. When the communication reads as automated, the inference is immediate and mostly subconscious: *if they automated the explaining, what else got automated?* The claims inherit the prose’s cheapness. An accurate finding wrapped in synthetic language reads less true than it is.

That inference is unfair to the many teams using AI drafting responsibly. It’s also how trust actually behaves, and pretending otherwise is how institutions get surprised.

Using AI in the pipeline without sounding like it

The workable position isn’t abstinence — it’s a pipeline where the model does translation and the scientist does truth and voice.

What AI does well here: first-pass jargon translation, length variants for different channels, plain-language analogy candidates. Journal authorship policies (Nature’s January 2023 position is the touchstone: an LLM can’t be an author) govern the research paper; communication materials are broadly fair game for disclosed assistance.

What the scientist must do, non-negotiably:

  1. Verify every simplified claim against the paper. Simplification is where overclaiming sneaks in — “associated with” becomes “causes,” “in mice” evaporates, effect sizes inflate into “breakthrough.” A model optimizes for readable, not for true-to-the-data. This check is the whole ballgame; a hyped summary does more damage than a boring one.
  2. Restore the surprise. “We expected the effect to vanish at low temperatures — it doubled” is the sentence journalists quote and readers remember, and it does not exist in your abstract.
  3. State the limitation like you mean it. “This is in cell cultures; a decade from any therapy” reads as integrity precisely because hype-machines never say it. Honest scope is the strongest human signal available.
  4. Say a real quote out loud. The AI-drafted “we are excited about the potential implications” quote is now a genre joke. Say the sentence to a colleague, then write down what you actually said — puzzlement, specificity, and all.
  5. Keep the process detail. The failed first approach, the instrument that had to be rebuilt, the Tuesday nights. Work has texture; abstracts don’t.

A practical register check before the release goes out: run it through a free AI detector with sentence-level output. Boilerplate paragraphs will flag — that’s the genre, ignore it. But if the *quote* or the finding explanation flags as machine-like, that’s a real finding: the paragraphs that must carry human signal are carrying none. Fix those by hand; for bulk-produced explainer content, a humanizer pass can help with register, but the verification and the voice have to come from the person who did the science — there’s a fuller workflow in how to humanize an AI-generated press release.

The short version scientists can tape above the monitor: let the model translate, never let it testify. The public can forgive dense writing. What it doesn’t forgive — at exactly the moment institutional trust is scarce — is discovering that nobody home wrote the words. More for research communicators on the blog.

Frequently Asked Questions

Why do science press summaries get flagged as AI-generated?

Press summaries live at the intersection of two formulaic registers: institutional press-release boilerplate (‘researchers at X University have discovered…’) and simplified science-explainer prose. Both are highly conventional, and conventional means statistically predictable — which is all a detector measures. Add that many summaries now genuinely are AI-drafted from the paper’s abstract, and the register has become doubly suspect. A flag on a press summary tells you it reads like every other press summary; whether a human or a model produced that sameness, the credibility cost with journalists and readers is similar.

Is it acceptable for scientists to use AI to draft lay summaries?

Widely, yes — translating a technical result into plain language is one of the better uses of a language model, and most journal and funder policies permit disclosed AI assistance in communication materials, though authorship policies like Nature’s bar crediting AI as an author on the research itself. Two disciplines keep it honest: the scientist verifies every simplified claim against what the paper actually shows, because simplification is where overclaiming sneaks in; and the final pass restores the human elements — why the result surprised you, what it doesn’t mean, what happens next in the lab.

Does it matter if the public can tell science writing is AI-generated?

It matters more than in almost any other field, because science communication’s entire product is trust. Readers who notice generated-sounding language in a research summary make a short inference: if the institution automated the explanation, how much scrutiny did the claims get? That inference may be unfair — an AI-drafted, scientist-verified summary can be perfectly accurate — but public trust in institutions doesn’t run on fairness. Recognizably synthetic prose in outreach materials spends credibility that took decades to accumulate, at a moment when science can least afford it.

What makes a research summary read as human-written?

The elements no model can supply from the abstract: the moment of surprise (‘we expected the effect to vanish at low temperatures — it doubled’), the honest limitation stated plainly, the detail of how the work actually felt (‘three years of Tuesday-night telescope time’), and a scientist speaking in first person about what puzzled them. Quotes help only if they’re real quotes — the AI-drafted ‘we are excited about the potential implications’ quote has become its own tell. Specificity about process and genuine uncertainty are the register of an actual human who did actual work.

Should press offices run releases through an AI detector before sending?

As a register check, it’s worth the two minutes. A sentence-level detector pass shows which paragraphs read as pure boilerplate — usually the opening and the institutional-pride paragraph — and those are exactly the ones journalists skim past anyway. The goal isn’t a low score; formulaic release structure will always score elevated. The goal is knowing which paragraphs carry zero human signal and fixing the ones where humanity matters: the finding explanation, the researcher quote, the limitations. If the quote paragraph flags as machine-like, that’s a real problem — it means your scientist’s voice isn’t in it.

Try it on your own text

Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.