PaperBleachPaperBleach
logo

How Top-K Token Rank Distributions Separate Human and Machine Writing

P
Paperbleach

07 Jul 2026

Here’s a small mental experiment. Take any word in a sentence, hide it, and ask a language model to rank every word that could plausibly come next, best guess first. The word you actually wrote lands somewhere on that list. First place? Fortieth? Eight-thousandth? That position is the whole idea behind one of the older, sturdier tricks in AI detection.

Detectors don’t look at one word, of course. They collect the rank of every word in a passage and study the *shape* of the pile. And it turns out that pile looks different depending on whether a person or a machine did the writing.

Key takeaways

  • Token rank is the position of each word in a model’s ranked list of likely next words. Rank 1 means the model would have guessed that exact word first.
  • Machine text tends to cluster in low ranks (top few guesses), because generators usually sample from the words the model already thinks are likely.
  • Human text spreads out. We reach for the odd word more often, so more of our words land in higher rank bands.
  • The classic tool for seeing this is GLTR, which colors each word by which rank band it falls into.
  • A low-rank distribution is a statistical signal, not proof. Clean, conventional human writing lands in low ranks too, which is where false positives come from.

Rank, not probability

Most explanations of AI detection start with perplexity, which is built from the exact probability a model assigns each word. Rank is the stripped-down cousin. Instead of keeping the probability, it keeps only the ordering.

Say the model is looking at “the coffee was too” and deciding what comes next. Internally it produces a probability for every word in its vocabulary. “Hot” might be 41%, “strong” 12%, “bitter” 7%, and so on down to words like “onomatopoeia” with a probability so small it’s basically zero. If you wrote “hot,” your token got rank 1. If you wrote “bitter,” rank 3. If you wrote “purple,” maybe rank 6,000.

Why bother throwing away the probability and keeping only the rank? Because rank is stubborn in a useful way. It doesn’t care whether the top guess had 92% probability or a shaky 30%. A word that was the model’s third choice is rank 3 either way. That stability makes rank less jumpy than raw probability across different topics and models, which is part of why rank-style features keep showing up inside detectors even as the fancier machinery around them changes.

Why machine text piles up low

Text generators don’t pick words at random. They sample from the model’s own probability distribution, and most sampling settings pull heavily toward the likely end. Top-k sampling literally restricts the choice to the k most probable words. Nucleus sampling keeps whatever words make up the top slice of probability mass. Even with some randomness dialed in, the output leans toward words the model ranked highly, because that’s the machinery producing it.

So when a detector later reads that text and ranks each word, it finds a suspicious amount of the writing sitting in the top handful of guesses. Word after word was something the model would have offered near the top of its own list. That’s not a coincidence; it’s a fingerprint of the process that made it.

Humans don’t work that way. We have a specific point to make, a memory to reach for, a joke we’ve been saving, a regional turn of phrase. We routinely pick the word that a language model would rank tenth or hundredth, not because we’re trying to be surprising but because we’re trying to say a particular thing. Over a paragraph, that scatters our words across the rank bands instead of hugging the bottom.

The four buckets GLTR uses

The tool that made this visible is GLTR, from a 2019 paper by Sebastian Gehrmann, Hendrik Strobelt, and Alexander Rush. It doesn’t hand you a percentage. It highlights every word in one of four colors based on which rank band it fell into: the top 10, the top 100, the top 1,000, or everything beyond that. Green for top 10, yellow, red, and purple for the rare tail.

Paste in raw GPT output and the screen lights up green with some yellow. Word after word was among the model’s top ten guesses. Paste in a messy human paragraph, a real one, with an aside and an unexpected verb, and you get a confetti of red and purple mixed through the green. You can see the color-coded token rank view GLTR made famous laid out in more detail, but the gist is that the *proportion of colors* is the tell, not any single word.

A worked example: two ways to end a sentence

Take the opening “After the storm passed, the street was.”

A generator, sampling toward the likely, is prone to finish with something like “quiet and covered in debris.” Every one of those words is a solid top-band guess. Ranked afterward, they come out green: 1, 2, 4, 3, 2. The distribution is a tight little cluster near the floor.

A person who was actually there might finish “shellacked, and weirdly loud with dripping.” “Shellacked” is a rank-2,000 word here, an odd but vivid choice. “Weirdly” as an adverb in that slot is high-rank too. Now the distribution has spikes: a couple of low ranks, then two words way out in the tail. Same sentence length, same grammar, completely different rank profile.

Neither passage is better writing by decree, though most readers feel the second one. The point is narrower: the *statistics* of the two are far apart, and a rank-based detector keys on exactly that gap.

The tempting, wrong lesson

You can see where someone might take this: if high ranks look human, just cram in high-rank words. Run the draft through a synonym tool, swap “hot” for “torrid,” “quiet” for “hushed,” and watch the ranks climb.

It doesn’t hold up. Human writing isn’t uniformly surprising. We don’t reach for the rare word every time; we mix the obvious and the unusual in a rhythm that fits the meaning. Force the rank up on every word and you produce a *new* unnatural distribution, oddly flat at the high end, plus prose that reads like it was translated by a dictionary. Detectors that weigh more than one feature notice the flatness, and human readers notice the strain. You’ve traded one detectable pattern for another and made the writing worse in the bargain. The reliable route is the same as always: mean what you say, and the odd words show up on their own in the places they belong. If you want to watch which parts of a draft read as low-rank and flat, you can run a draft through the free checker and edit from there.

What the rank distribution can’t tell you

Here’s the limit worth burning into memory. A low-rank distribution says the words in a passage were, on the whole, predictable. It does *not* say a machine wrote them.

Clear writing is often predictable on purpose. Instructions, safety notices, tidy expository prose, the careful English of a non-native writer who learned the conventional forms first, all of it tends to use the expected word because the expected word is the clearest one. That writing lands in low ranks and can trip a rank-based flag with no AI anywhere in the story. This isn’t a rare edge case. A 2023 Stanford study in *Patterns* found detectors flagged more than half of TOEFL essays by non-native English speakers as AI, largely because conventional, lower-surprise phrasing reads as machine-like to these statistics.

And the whole enterprise is an estimate sitting on a threshold, not a verdict. OpenAI, which builds the generators, pulled its own text classifier in July 2023 over low accuracy. If the company with the most to gain from reliable detection couldn’t get there, treat any rank-based reading, GLTR’s colors included, as one piece of evidence about statistical patterns rather than a statement about who held the pen. For the fuller picture of how these token-level numbers fit together, we walk through log-likelihood, rank, and entropy as detection statistics separately.

Frequently asked questions

What does “token rank” mean in AI detection?

For each word, a model can rank every possible next word from most to least likely. The word you actually used sits at some position in that list, its rank. A detector collects the rank of every word and studies the whole distribution. Machine text piles up in the low ranks because generators tend to pick words the model already rated highly.

How is token rank different from perplexity?

Perplexity uses the exact probability of each word; rank keeps only the ordering. Rank is coarser but sturdier, because it ignores whether the top guess was 60% or 92% likely. A word that was the model’s third choice is rank 3 regardless, and that stability is why rank features survive across topics and models.

Can I lower my token ranks on purpose to look human?

You can, but it usually backfires. Swapping in rare words everywhere is thesaurus-humanizing, and it leaves a flat, high-rank distribution that’s its own tell, plus stilted prose. Real writing mixes obvious and unusual words in a natural rhythm, which is hard to fake by brute force.

Does a low-rank distribution prove I used AI?

No. Clear, conventional human writing lands in low ranks too, so the signal produces false positives on exactly the writing that plays it straight. It’s a probability about patterns, not evidence of authorship.

Which detectors use token rank?

The clearest example is GLTR, which color-codes words by rank band. Rank also appears as a feature inside broader detectors and in zero-shot methods comparing log-probability against log-rank. Most commercial tools don’t reveal their exact features, but rank statistics are a long-standing part of the field.

The short version

Rank asks a simple question of every word: how high on the model’s own guess list did you land? Machine text answers “near the top, over and over.” Human text answers with a scatter, plenty of top guesses, but also the occasional word from deep in the tail. That difference in shape is real and it’s old, which is why it keeps showing up inside newer tools. It’s still a statistic on a threshold, though, and the writing that plays it straight, clear, conventional, careful, can look just as low-rank as a machine. If a draft you wrote reads as flat, run a draft through the free checker, then add back the specific words the meaning was already asking for. You can also browse more on how detection works or see the plans.

Try it on your own text

Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.