PaperBleachPaperBleach
logo

How Sentence Embeddings Capture Burstiness That Word Counts Miss

P
Paperbleach

06 Jul 2026

Most explanations of burstiness stop at sentence length. Humans mix short punchy lines with long winding ones, the story goes, while machines hum along at an even clip. True, as far as it goes. But length is only the surface of variation. There’s a deeper kind that a ruler can’t measure — how much the *meaning* jumps around from one sentence to the next — and catching it takes a different tool entirely.

That tool is the sentence embedding. It’s how detection is starting to read a signal that word counts are simply blind to: not how long your sentences are, but how far apart they are in meaning.

Key takeaways

  • Classic burstiness measures surface variation — sentence length, per-sentence perplexity — and misses variation in meaning.
  • A sentence embedding turns a whole sentence into a point in space that captures its meaning, not its exact words.
  • Measuring how far apart consecutive sentences sit in that space gives you *semantic* burstiness.
  • Human writing tends to wander in meaning; first-pass model output tends to march straight down the topic.
  • It’s a real, complementary signal — and, like every signal, useless as a standalone verdict.

What plain burstiness can and can’t see

Burstiness is the second big statistic detectors lean on, alongside perplexity — we cover both in the statistics behind perplexity and burstiness, and their differences in burstiness versus perplexity as separate signals. The usual version measures how much your surface statistics bounce: sentence lengths, or the predictability score from sentence to sentence. High variance reads human, flat reads machine.

But surface variation has a blind spot, and it’s a big one. You can vary your sentence lengths beautifully while never leaving your topic — every sentence a different length, all of them saying closely related things. And you can hold your sentence lengths perfectly even while leaping between loosely connected ideas. Length tells you nothing about *meaning*. A writer can be metrically bursty and semantically monotonous, or the reverse. Word-count burstiness measures the rhythm of the prose and stays deaf to the rhythm of the thought.

Embeddings turn meaning into geometry

To measure the rhythm of the thought, you need a way to compare what sentences *mean*, and that’s exactly what a sentence embedding provides.

An embedding takes a full sentence and maps it to a point in a high-dimensional space — a long list of numbers — arranged so that meaning becomes distance. “The dog chased the ball” and “A puppy ran after the toy” share almost no words, but they mean nearly the same thing, so their points land close together. “The dog chased the ball” and “Interest rates rose in the third quarter” land far apart. The technology behind this is well established; approaches like Sentence-BERT, introduced by Reimers and Gurevych in 2019, made it practical to embed sentences so that semantic similarity shows up as geometric closeness.

Once every sentence in a passage is a point, a document stops being a string of words and becomes a *trajectory* — a path hopping from point to point as the meaning moves. And a trajectory is something you can measure.

Measuring the jumps

The standard yardstick for “how close are two embedding points” is cosine similarity, which scores how aligned two vectors are: near 1 when two sentences mean almost the same thing, near 0 when they’re unrelated. Take each pair of neighboring sentences in a piece, measure their cosine similarity, and you get a sequence of numbers describing how much the meaning shifts at every step.

Now the pattern in those numbers is the whole point. If consecutive sentences are always highly similar, the writing is gliding along smooth rails — every sentence a small, safe step from the last. If the similarities swing between high and low, the meaning is lurching: a tight cluster of related sentences, then a jump to something new, then a digression, then back. That spread — the variance in how much meaning moves sentence to sentence — is semantic burstiness. High spread, bursty thought. Low spread, monotone thought. And crucially, none of this depends on how long the sentences are.

Why humans jump more

Here’s the behavioral observation that makes the signal worth measuring. People don’t think in a straight line, and it shows in their writing. Real prose takes tangents — a sudden aside, a personal example, a loosely related point that occurred to the writer mid-paragraph, a doubling-back to qualify something. Those moves create genuine jumps in the meaning-trajectory, moments where consecutive sentences sit far apart in embedding space.

First-pass model output tends to behave differently. Ask a model to write about a topic and it usually stays admirably, relentlessly on that topic, with smooth transitions engineered between each sentence and the next. The result is high semantic coherence and low semantic burstiness — a trajectory that inches along without ever really jumping. It’s not a hard rule; models can be prompted into digression and humans can write with laser focus. But as a tendency, the difference is real, and embeddings surface it where sentence-length variance saw nothing at all.

Think of it as the difference between a guided tour that hits every stop in order and a friend showing you around who keeps wandering off to point at something. The tour is coherent. The friend is bursty. Word counts can’t tell them apart; the meaning-trajectory can.

The honest limits

Everything good about this signal comes with the same caveat that haunts every detection feature, and it deserves stating plainly.

Plenty of excellent human writing is low in semantic burstiness. Technical documentation stays rigorously on-topic by design. A well-disciplined formal report doesn’t digress. A focused, tightly-argued essay may hold a steady semantic line from start to finish precisely because it’s *good*. Lean on semantic burstiness alone and you’d flag exactly the careful, focused writers you least want to accuse — the same pattern of collateral damage that plagues perplexity-based detection.

So the right way to hold this: semantic burstiness is a complementary lens, not a replacement. It adds a meaning-level view on top of the surface statistics, which is genuinely valuable because it catches variation the surface hides. Combined with other signals it can strengthen a detector’s read. On its own, treated as proof, it fails the same way everything else does. The lesson repeats across the whole field — more signals can sharpen the picture, and none of them turns a probability into a fact.

Frequently asked questions

What’s the difference between word-count burstiness and semantic burstiness?

Word-count burstiness looks at surface variation — how much sentence lengths or perplexity scores bounce. Semantic burstiness looks at variation in meaning: how much the topic and focus shift sentence to sentence. You can write with lots of length variation while staying rigidly on-topic, or with even lengths while wandering. Embeddings measure the second kind, which surface counts can’t see.

What is a sentence embedding, in simple terms?

It’s a way of turning a whole sentence into a list of numbers — a point in space — that captures its meaning rather than its exact words. Two sentences that say similar things land close together even with no shared vocabulary; two that say different things land far apart. Once every sentence is a point, you can measure how far apart consecutive sentences are.

Why would human writing have more semantic burstiness than AI?

Because people digress. Real writing takes tangents, doubles back, drops in an aside, jumps to a loosely related point, then recovers — producing larger meaning-jumps between sentences. First-pass model output tends to march straight down the topic with smooth transitions: high coherence, low semantic burstiness. The embeddings expose that steadiness as a soft signal.

How do embeddings actually measure the variation?

Usually with cosine similarity, which scores how aligned two embedding vectors are. You measure the similarity of each neighboring pair of sentences and look at the pattern across the piece. Consistently high similarity means smooth, on-rails writing; a mix of high and low means the meaning is jumping around. The spread of those similarities is the semantic burstiness.

Is embedding-based detection reliable on its own?

No single signal is. Plenty of good human writing is tightly focused and low in semantic burstiness — technical docs, formal reports, disciplined essays — so relying on it alone would falsely flag careful writers. It’s best understood as one more feature that adds a meaning-level view to the surface statistics: useful in combination, never a verdict by itself.

The short version

Burstiness isn’t only about sentence length. There’s a deeper variation — how much the meaning jumps from sentence to sentence — that surface counts can’t touch, and sentence embeddings measure it by turning each sentence into a point and watching how far the writing travels between them. Human thought wanders; first-pass model output tends to glide. It’s a smart, complementary signal that catches what word counts miss, and it still isn’t proof, because focused human writers are semantically smooth too. Want to see how your own writing reads? Run a draft through the free checker, then browse more detection guides or see the plans.

Try it on your own text

Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.