PaperBleachPaperBleach
logo

DetectLLM and the LRR Metric: A Newer Spin on Zero-Shot Detection

P
Paperbleach

06 Jul 2026

DetectGPT proved you could catch AI text without training data — just poke the writing, watch the probabilities move, and read the curvature. Clever, but expensive: scoring one passage meant generating and re-scoring dozens of reworded copies. A 2023 method called DetectLLM asked a sharp follow-up question. What if you could get a comparably strong signal in a *single pass*, no perturbations at all, by using a statistic DetectGPT largely ignored?

That statistic is rank. And the metric built on it, LRR, is a nice example of how a small change in what you measure can buy a lot of efficiency.

Key takeaways

  • DetectLLM, from Su and colleagues in 2023, introduced two zero-shot detection metrics: LRR and NPR.
  • LRR is the Log-Likelihood Log-Rank Ratio — it combines how probable your words were with how highly ranked they were, and needs no perturbations.
  • The ratio amplifies the machine-text signature: predictable words that also sit near the top of the model’s ranking.
  • NPR is a perturbation-based cousin, more accurate but slower, that watches how log-rank degrades when text is disturbed.
  • It’s a sharper, cheaper signal — not a solution to detection’s underlying reliability problems.

The gap DetectLLM set out to fill

To place DetectLLM, start with what came before. Trained classifiers learn from labeled examples and go stale when new models arrive. Zero-shot methods skip training and read a scoring model’s probabilities directly. DetectGPT’s curvature trick was the standout zero-shot approach — accurate, but it paid for that accuracy with heavy computation, because measuring curvature meant perturbing the text many times over.

DetectLLM’s authors noticed something the field had underused: the *rank* of each token, not just its probability. Older tools like the GLTR visualizer had shown that rank information is revealing, but the newer zero-shot detectors leaned mostly on likelihood. DetectLLM’s bet was that folding rank back in could produce a stronger signal, cheaply.

Likelihood and rank: two views of the same word

To follow LRR you need two token statistics, and they’re closely related but not the same. For a fuller tour, we cover log-likelihood, rank, and entropy as token statistics separately, but here’s the compact version.

At each position, a language model produces a full ranked list of possible next words with probabilities attached.

  • Log-likelihood answers: *how probable was the word you actually used?* Higher (less negative) means the model found your word likely.
  • Log-rank answers: *where did your word sit in the ranking?* If your word was the model’s #1 pick, its rank is 1; if it was the 500th most likely option, its rank is 500. Log-rank is small when your word ranked near the top.

Why keep both? Because they can disagree in useful ways. Two words might have similar probabilities yet very different ranks — a word can be “reasonably likely” while still being outranked by many others, or “the clear favorite” in a spot where the model was confident. Rank carries ordering information that raw probability smears over. Machine text, which tends to pick top-ranked words, leaves a fingerprint in the ranks that likelihood alone can blur.

What LRR actually does

LRR — the Log-Likelihood Log-Rank Ratio — puts these two together as a ratio rather than using either alone. Intuitively, it divides a measure of how *probable* the text was by a measure of how *highly ranked* it was, and that combination sharpens the contrast between human and machine writing.

Here’s the intuition for why a ratio helps. Machine-generated text is doubly predictable: its words are both high-probability *and* high-rank (near the top of the list). Human text is lumpier on both counts — the occasional surprising word is both less probable and more deeply buried in the ranking. By taking the ratio, LRR lets the two signals reinforce each other. When both point the same way, the ratio swings hard, giving a crisper separation than either statistic manages on its own. It’s a bit like judging a student not just by their grade but by their class rank — the two together tell you more than one alone.

The payoff is efficiency. Computing LRR takes one forward pass to get the probabilities and ranks. No perturbing, no re-scoring fifty variants. That makes it dramatically cheaper than curvature-based detection while, in the DetectLLM experiments, holding up competitively — sometimes better — on the text they tested.

NPR: the accurate, slower sibling

DetectLLM didn’t stop at LRR. It also introduced NPR, the Normalized Perturbed log-Rank, for cases where you’ll pay more compute for more accuracy.

NPR brings back perturbation, DetectGPT-style — disturb the text, re-score it — but instead of watching how the log-*likelihood* moves, it watches how the log-*rank* moves. The insight is that machine text’s rank degrades more sharply when you perturb it: nudge a passage that was sitting on top-ranked words and those words tend to tumble down the ranking, while human text’s already-scattered ranks move less predictably. NPR captures that by comparing the perturbed rank to the original. It’s more accurate than LRR in their results, at the cost of the expensive perturbation step LRR was designed to avoid. So DetectLLM effectively offers a dial: LRR when you want speed, NPR when you want the extra accuracy and can afford the compute.

Where it still runs into the same walls

DetectLLM is a real, well-motivated improvement. It’s also not a way out of the problems that dog every detector, and it’s worth being clear-eyed about that.

It’s still a white-box, model-relative method. LRR and NPR need access to a scoring model’s probabilities and ranks, and like DetectGPT they work best when that scoring model resembles the one that actually wrote the text. Point them at output from an unknown or very different model and the signal softens — the same cross-model penalty the whole zero-shot family shares.

It still outputs a statistical judgment, not proof. A cleaner ratio is still a ratio, and it can be wrong. Plain, careful, or non-native writing tends to be low-perplexity and high-rank for the same reasons machine text is, so the familiar false positives don’t vanish just because the math got tidier. And it inherits the field’s broader fragility — paraphrasing and editing that scramble the token statistics will scramble LRR and NPR right along with them.

So the honest framing is that DetectLLM sharpened the tool without changing what the tool fundamentally is. That’s genuinely valuable — cheaper, crisper detection is worth having — but it’s an upgrade to the instrument, not a resolution of the questions the instrument can’t answer.

Frequently asked questions

What is the LRR metric in DetectLLM?

LRR stands for Log-Likelihood Log-Rank Ratio. It combines how probable the model found each word (log-likelihood) with how highly it ranked each word (log-rank) into a single ratio. Machine text tends to score high on both — probable words that also sit near the top of the ranking — and the ratio amplifies that into a cleaner signal than either number alone. It’s DetectLLM’s fast, single-pass detection method.

How is DetectLLM different from DetectGPT?

DetectGPT perturbs your text dozens of times and re-scores each version — accurate but slow. DetectLLM’s LRR needs no perturbations; it reads log-likelihood and log-rank in a single pass, so it’s far cheaper. DetectLLM also offers NPR, a perturbation-based metric that trades speed for accuracy, but the headline idea is a strong signal without all the rewriting.

What is log-rank and why does it help?

For each word you wrote, the model ranks every possible word by probability, and log-rank captures where your word landed. If your words are consistently the model’s top pick, the ranks are tiny — a machine-text tell. Log-likelihood measures probability, but log-rank adds ordering it misses: two words can have similar probabilities yet very different ranks, and DetectLLM uses that extra angle.

What is the NPR metric?

NPR is DetectLLM’s Normalized Perturbed log-Rank. Like DetectGPT it perturbs the text and re-scores, but it tracks how the log-rank changes rather than the log-likelihood. Machine text’s rank degrades more sharply under perturbation, so the ratio of perturbed to original rank separates AI from human text. NPR is usually more accurate than LRR but slower, since it revives the expensive perturbation step.

Does DetectLLM solve the reliability problems of AI detection?

No. It’s a cleaner, more efficient signal, not a cure. It still assumes access to a scoring model’s probabilities, still works best when that model resembles the one that wrote the text, and still outputs a statistical judgment that can be wrong — including false positives on plain or non-native writing. A genuine improvement in the toolkit, not an escape from the field’s limits.

The short version

DetectLLM’s contribution is rank. By reading not just how probable your words were but how highly the model ranked them, its LRR metric squeezes a strong zero-shot signal out of a single pass — no perturbation marathon required — and its NPR metric offers extra accuracy when you’ll pay for the compute. It’s a smart, efficient refinement of the DetectGPT idea. It’s also still model-relative, still statistical, and still capable of flagging honest writing, because the underlying limits didn’t move. Curious how your writing scores? Run a draft through the free checker, then browse more detection guides or see the plans.

Try it on your own text

Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.