What an AI Detector’s ‘Perplexity Per Word’ Chart Is Showing You
07 Jul 2026
Some detectors don’t just hand you a percentage. They show you a chart: a jagged line or a row of bars running across your text, one mark per word, rising and falling. It looks technical and a little intimidating, and most people glance at it, decide it’s proof of something, and move on. It’s actually the most honest thing these tools show you, more honest than the score, if you know how to read it.
The whole chart is answering one small question, over and over, once per word: *how surprised was the model to see this?*
Key takeaways
- The chart plots perplexity per word: how surprised a language model was by each word, given the words before it.
- Tall marks mean the word was unexpected; short marks mean it was the expected choice.
- Low and flat tends to read machine-like (consistently predictable); tall and jagged tends to read human (varied, surprising in places).
- A low, flat chart does not prove AI. Clear, conventional, careful writing looks flat too, which is where false positives come from.
- The chart’s best use is as an editing map: flat stretches are your most uniform, most generic passages.
One question, asked per word
Start with what “perplexity” means at the level of a single word, because the chart is just that idea repeated. Take the phrase “she poured the coffee into her.” A language model, reading up to “her,” has strong expectations for the next word. “Cup” or “mug” is barely surprising; the model would have guessed it. So if your next word is “cup,” that spot on the chart is *short*. Low surprise.
Now suppose the next word is “shoe.” The model did not see that coming. Surprise spikes, and that spot on the chart is *tall*. Same sentence up to that point, wildly different surprise for the final word.
That’s the entire chart. It walks through your text and, at each word, marks how far the actual word was from what the model expected. Left to right, you get a landscape of surprise: valleys where you wrote the predictable word, peaks where you wrote the unexpected one. If you want the concept built up slowly from scratch, we have the plain-English walkthrough of what perplexity is. Here we’re just learning to read the picture.
Reading the two shapes
Two features of the landscape carry almost all the meaning: how *high* it sits on average, and how *jagged* it is.
Height is overall predictability. A chart that hugs the floor means word after word was the expected choice, low perplexity, the pattern machine text tends to produce because generators lean toward likely words. A chart that rides higher means you reached for the unexpected more often.
Jaggedness is variation, the thing sometimes called burstiness. A line that spikes up and drops down at irregular intervals means your surprise is uneven, some very predictable words, some very surprising ones, which is how humans naturally write. A line that stays at one steady level, high or low, means uniform predictability, which is more machine-like.
Put them together and you get the rough reading everyone uses:
- Low and flat: consistently predictable, little variation. Reads machine-like.
- Tall and jagged: often surprising, highly variable. Reads human.
- Low but jagged, or high but flat: mixed signals, the interesting middle, where a synonym-swapped AI draft (raised height, still flat) or a plain-spoken human (low height, but jagged rhythm) confuses the simple story.
This is the same information GLTR renders as colors instead of a line, greens for the expected words, reds and purples for the surprising ones, and the 2019 GLTR paper by Gehrmann, Strobelt, and Rush is where the “look at the per-word landscape” idea got its clearest public form. A token-level heatmap tells you about your draft in exactly this spirit; the chart is the same signal drawn as a graph.
A mini-scenario: two charts, one topic
Picture two paragraphs about a rainstorm, charted side by side.
The first is generic AI output: “The rain fell steadily throughout the afternoon, and the streets became wet and difficult to navigate.” Every word is a solid expected choice. Its chart is a low, calm ripple near the floor, no real peaks, no deep valleys. Flat and low.
The second is something a person actually saw: “Rain hammered down all afternoon, and by four the gutters were choking on cigarette butts and one lost sandal.” “Hammered,” “choking,” “cigarette butts,” “sandal”, each is a spike, an unexpected but exact word. Its chart is a mountain range: low stretches for the connective tissue, sharp peaks at the specifics. Tall and jagged.
Same length, same subject. The charts don’t look remotely alike, and the difference is entirely in word choice and specificity. That’s what the picture is capturing, and it’s why the second paragraph reads as human to both a detector and a person.
The trap: flat doesn’t mean machine
Here’s the part that has to travel with everything above, or the chart becomes dangerous. A low, flat line means your writing was *predictable*. It does not mean a machine wrote it.
Predictability has plenty of innocent sources. A safety instruction is flat because clarity demands the plain word. A legal or technical passage is flat because precision leaves little room to be surprising. A definition is flat because there’s really one clear way to say it. And a careful non-native English writer is often flat because they learned the conventional, correct forms and use them faithfully. All of that produces a low, calm chart, and none of it involves AI.
This isn’t a minor caveat; it’s the central failure of reading the chart as proof. The same 2023 research that found detectors misclassifying most non-native English essays is, in chart terms, a story about flat lines being mistaken for machines. The picture shows predictability accurately. The leap from “predictable” to “machine-written” is the error, and it’s the one that gets honest people accused. Even OpenAI’s own classifier, built by the people who make the models, was retired in 2023 for low accuracy, because “looks predictable” is simply not the same as “was generated.”
The chart’s real value: it’s a map
So what’s the chart *good* for? Editing. It’s the most useful diagnostic these tools offer, precisely because it shows detail the percentage hides.
Wherever the line runs low and flat for a long stretch, you’ve found a passage that reads as uniform, and that passage is almost always a bit generic to a human reader too. That’s not a coincidence; flatness and blandness are the same thing measured two ways. So the chart hands you a to-do list: go to the flat valleys and lift them. Add a concrete specific. Use the exact odd word instead of the safe one. Break a long even sentence with a short sharp one. Each of those changes pokes the line up and roughens it right where it was dead, and, not incidentally, makes the writing better.
That’s the honest version of “humanizing,” and it’s the opposite of gaming. You’re not tricking the chart; you’re responding to what it correctly noticed, that a stretch of your writing was uniform and predictable. To get this map for something you’re working on, run a draft through the free checker and treat the flattest passages as the ones worth your attention, whether or not a machine was ever near them.
Frequently asked questions
What does a perplexity-per-word chart actually plot?
For each word, how surprised a model was to see it given the prior words. Tall means unexpected; short means the expected choice. Reading left to right, you watch surprise rise and fall word by word. Overall height shows how predictable the writing is; jaggedness shows how much that predictability varies.
What shape suggests machine writing?
A low, flat line. Generators favor likely words, so machine text sits at consistently low surprise and doesn’t spike much. Human writing is usually taller and jagged, because we reach for the unexpected, specific word at irregular intervals. Flat-and-low reads machine-like; tall-and-spiky reads human, though these are tendencies, not rules.
Does a low, flat chart prove I used AI?
No. Clear, conventional, careful writing produces a flat chart too, because clarity means using the expected word. Technical writing, instructions, and non-native English can all look flat with no machine involved. The chart shows predictability, which has innocent causes; reading it as proof is what produces false accusations.
How is the chart different from the single percentage?
The percentage is the whole chart crushed into one averaged, thresholded number. The chart shows the raw per-word detail, which words spiked, which sat flat, before that detail was thrown away. That makes the chart more informative than the score, and genuinely useful as an editing map.
Can I use the chart to improve my writing?
Yes, that’s its best use. Long low-flat stretches are passages a detector reads as uniform and that also tend to read as generic. Adding a concrete specific, an exact unexpected word, or a change of sentence length lifts and roughens the line there. You’re not gaming anything, you’re making genuinely less uniform writing.
The short version
A perplexity-per-word chart is a picture of one small question asked at every word: how surprised was the model? Valleys are your predictable words, peaks are your surprising ones, and the overall shape, low and flat versus tall and jagged, is what a detector reads as machine-like or human. Just remember the shape shows *predictability*, not authorship: flat writing has innocent causes, and clear, careful, non-native prose lives in the valleys through no fault of anyone’s. Read the chart as an editing map, lift your flattest passages with real specifics, and never mistake a low line for a confession. To see the chart for your own draft, run it through the free checker, browse more on how detection works, or see the plans.
Try it on your own text
Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.


