How AI Detectors Score Bullet Points, Headers, and Markdown Formatting
06 Jul 2026
Paste a clean, well-structured document into an AI detector — the kind with tidy headers, a few bullet lists, some bolded key terms — and you can get a wobblier result than you’d get from the same content written as flat paragraphs. That’s not your imagination, and it’s not a sign the formatted version is “more AI.” It’s a sign that formatting does something to the machinery underneath, and the machinery wasn’t really built for it.
Detectors were trained on prose. Markdown is not prose. What happens in the gap between those two facts is worth understanding, because it explains a lot of otherwise baffling scores.
Key takeaways
- Detectors expect flowing sentences. Headers, bullets, and markup are not that, so they get handled less reliably.
- Whether the raw symbols (
##,**,-) affect your score depends on whether the tool strips them before scoring, and you usually can’t tell. - Headers are short, oddly-shaped fragments that produce noisy, low-confidence chunk scores.
- Reformatting can move the number, which is a sign the number was unstable, not that your authorship changed.
- The same blind spot cuts both ways — it causes false positives on human writing and can let AI text slip through.
Detectors read prose, and markdown isn’t prose
Almost every AI detector was built around a simple assumption: the input is running text, sentences following sentences, long enough to measure how predictable the words are. That assumption is baked deep into the pipeline. The tool breaks your input into sentences, scores each one for predictability, and combines those scores. As GPTZero and others describe, the whole thing runs on measuring predictability and its variation across ordinary written language.
Markdown quietly violates the assumption. A header isn’t a sentence. A bullet isn’t a sentence. A bolded phrase in the middle of a line carries invisible symbols the model has to reckon with. None of this is what the detector rehearsed on, so it improvises — and improvisation is where the noise creeps in.
The symbols: stripped, or scored?
The first fork in the road is what the tool does with the raw markup characters. There are two camps, and they behave completely differently.
Strippers. Some detectors preprocess your text first, pulling out markdown syntax so only the human-readable content reaches the scoring model. In this camp, ## Key Findings becomes just Key Findings, and the ## never influences anything. Cleaner, but you’re now at the mercy of how aggressively they strip — some also drop the header text entirely, some keep it as a floating fragment.
Passers. Other detectors take whatever you paste, symbols and all, and feed it straight to the model. Now important includes those asterisks as actual tokens. To a language model, a double asterisk mid-word is an unusual, low-content token, and stringing several markup symbols together produces sequences that don’t look like natural language at all. That can tug the predictability statistics in directions nobody intended — sometimes making text look more surprising, sometimes less.
The maddening part: you’re rarely told which camp a given tool falls into. So the same document can score differently across tools partly because of a preprocessing decision you can’t see.
Headers are tiny, weird sentences
Set the symbols aside and consider the header text itself. Headers are written in a clipped, contextless register: *Key Findings. Next Steps. How It Works. Background.* They’re grammatically nothing like the sentences around them — often just a noun phrase, two or three words, no verb.
For a detector that scores text chunk by chunk, each header is a minuscule fragment with almost no context to work from. And short fragments are exactly where predictability measures fall apart, because there isn’t enough text for the statistics to settle. One three-word header can score as wildly human or wildly AI on essentially a coin flip’s worth of evidence. Stack up a dozen headers in a long document and you’ve handed the tool a dozen noisy, low-information readings that get folded into the final number. This is a close cousin of the problem in how one flagged line skews a whole document — small, unreliable chunks distorting an aggregate.
Bullets bring their own trouble
Bullet points compound the effect from a different angle, and it’s the same core issue we cover in why lists and tables throw detectors off. Bullets tend to be short, parallel, and grammatically uniform: *Cut costs. Improve retention. Reduce risk.* That parallelism reads as low-variation, low-perplexity text regardless of who wrote it, because the list format itself flattens everyone’s phrasing into the same clipped shape.
So a heavily bulleted document gives a detector two problems at once — fragments too short to score confidently, and a uniformity that mimics the statistical signature detectors associate with machine writing. Neither has anything to do with whether a person or a model actually produced the content. The format is doing the talking.
Why the number moves when you reformat
Here’s the observation that ties it together, and the one worth internalizing. If you take the same content and toggle it between a bulleted, header-heavy layout and flat paragraphs, the AI score can shift noticeably. People sometimes read that as “the plain version is more human.” It isn’t. Your authorship didn’t change; only the formatting did.
What the swing actually reveals is that the score was fragile to begin with. Flowing paragraphs give the detector long, sentence-shaped input — the thing it was built to measure — so it produces a steadier reading. Chopped-up formatting gives it short, strange fragments and stray symbols, so it produces a jumpier one. A number that jumps when only the layout changed was never a solid measurement of anything about the writing. Treat that volatility as a warning label, not a target to optimize against.
The blind spot cuts both ways
It’s tempting to read all this as “add formatting to dodge detection,” and the symmetry is real: whatever makes human formatted text falsely flag can also let AI-generated formatted text slip through. A detector confused by markdown is confused in both directions.
But leaning on that is a poor plan. Preprocessing changes between tools and over time, so a gap that exists today may close tomorrow. And any reviewer who suspects formatting games will simply strip the layout and re-run the check on the plain text. The durable takeaway isn’t a trick — it’s a caution. When structured content gets a surprising score, the format is a prime suspect. Check the prose version before you conclude anything about the writer.
Frequently asked questions
Do the markdown symbols themselves, like ## or **, change my AI score?
It depends on whether the tool strips them first. Many detectors remove markup before scoring, so the symbols never reach the model. Others paste your raw text straight in, and then those symbols become unusual, low-content tokens that can nudge the statistics unpredictably. You often can’t tell which kind of tool you’re using, which is exactly the problem.
Why do my headers get treated oddly by detectors?
Headers are short, dense, and grammatically unlike sentences — “Key Findings,” “Next Steps.” A detector that scores text sentence by sentence sees each header as a tiny fragment with almost no context. Short fragments produce unstable, noisy scores, so a header-heavy document hands the tool many low-information chunks that pull the overall number around.
Does converting formatted content to plain paragraphs change the result?
It can, because paragraphs give the detector flowing sentences with enough length to measure predictability reliably. Flattening headers and bullets into prose usually produces a steadier reading. But don’t reformat just to move a number — the right format serves your reader. A score that swings on formatting alone is telling you the score was shaky, not that your authorship changed.
Are lists and tables handled differently from headers and bold text?
Somewhat. Lists and tables break text into short, uniform fragments, starving the detector of sentence-level variation. Inline formatting like bold or links is more about stray tokens surviving preprocessing. Both fall under the same umbrella: detectors were trained on ordinary prose, and anything that isn’t gets handled less reliably, in tool-specific ways.
Can formatting be used to dodge a detector on purpose?
In principle the same blind spots that cause false positives can cause false negatives, so heavy formatting can make AI text harder to score. But that’s fragile — tools change their preprocessing, and a suspicious reviewer will just reformat and re-check. It’s not durable, and it’s a bad reason to mangle a document’s structure.
The short version
AI detectors were built for flowing prose, and markdown isn’t that. Whether the raw symbols matter depends on invisible preprocessing you can’t inspect; headers arrive as tiny, contextless fragments that score noisily; bullets add a uniformity that mimics machine writing. The tell is that reformatting the same content can move the number — proof the reading was fragile, not proof of anything about the author. So when structured text gets a strange score, suspect the format first. Want to see how your own draft reads? Run it through the free checker, then browse more detection guides or see the plans.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.


