logo

Why the Same Essay Gets a Different AI Score in Canvas vs Standalone Turnitin

Paperbleach
Paperbleach

03 Aug 2026

The detector is identical in both places — when the same essay gets a different AI score in Canvas versus standalone Turnitin, the difference comes from what changed around the analysis: the extracted text, the timing, the settings, or the model version, not the venue. Finding your Turnitin AI score different in Canvas than in a direct check feels like catching the system contradicting itself. It’s actually a tour of everything that sits between “my essay” and “the number,” and once you can name those layers, the mystery mostly evaporates.

Key takeaways

  • Canvas doesn’t score anything. It’s an LTI pipe to the same Turnitin engine that handles direct submissions — there’s one detector, not two.
  • The usual culprits for a gap: different file formats extracting different text, submissions weeks apart across model updates, different assignment settings, or probabilistic wobble on borderline prose.
  • The AI indicator is an estimate with acknowledged uncertainty, especially at lower percentages — small variation is expected behavior, not malfunction.
  • The institutionally real score is the one on the actual assignment submission.
  • If scores diverge sharply, compare inputs first: file type, word count, what text actually got extracted.

One engine, two doors

Start by killing the premise. There is no “Canvas AI detector.” When your instructor attaches Turnitin to a Canvas assignment, your upload travels through the LTI integration to Turnitin’s servers, where the same models that process direct turnitin.com submissions process yours. Canvas contributes the door, not the judgment.

So a score gap between “in Canvas” and “standalone” is never the platform scoring differently. It’s the same judge reading two subtly different case files. The interesting question becomes: what made the files differ?

Culprit one: the text that actually got extracted

The detector doesn’t analyze your document — it analyzes text *extracted from* your document, and extraction is where identical-looking essays quietly diverge.

Submit a .docx one time and a PDF another, and the engine may receive meaningfully different text. PDF extraction can mangle line endings, merge or split words, drop headers, or — with image-based PDFs — recover nothing at all. Even the same PDF exported from different tools can extract differently. We’ve dug into how file type alone can move a Turnitin AI score, and it’s a bigger lever than most people expect.

There’s also composition: the AI indicator only analyzes qualifying prose. Quotes, reference lists, bullet fragments, and non-English passages get excluded. If one submission included your bibliography and another didn’t — or one route’s extraction handled your block quotes differently — the *denominator* changed. Same essay, different analyzed text, different percentage. In the extreme case, one version falls under the minimum word threshold and gets no score at all; we’ve covered why a Turnitin AI score can be missing in Canvas entirely.

Culprit two: time and model versions

Detection models aren’t frozen. Turnitin updates its AI classifier as generation models evolve, which means the detector that scores your essay in September is not bit-for-bit the detector from June. Students comparing a summer self-check against a fall assignment submission are comparing two different model versions and shouldn’t expect identical output.

This is the least intuitive cause, because nothing on your side changed. But it’s the nature of the arms race: a static detector decays as new generators appear, so vendors retrain, and retraining moves boundary cases. An essay whose prose sits comfortably in “reads human” territory barely notices. An essay near the decision boundary can swing visibly.

Culprit three: the estimate is probabilistic

Even holding the model constant, the AI indicator is a statistical estimate, not a deterministic measurement like word count. Turnitin itself flags this — lower-range scores carry an explicit caution because the confidence there is weaker. Text near the classifier’s boundary is exactly the text that can land on different sides of it across runs.

The mental model that serves you best: the score is a reading from an instrument with error bars, and the error bars are widest in the low-and-middle range. A shift from 8% to 14% between two checks is well within instrument noise. A shift from 5% to 70% is not — that points back at culprits one and four, because models and randomness don’t move honest prose that far.

Culprit four: assignment settings and accounts

Two Turnitin-backed checks are rarely configured identically. The Canvas assignment your instructor built carries institution- and assignment-level settings; a standalone or self-check submission carries different ones. Settings that exclude quotes or bibliographies change the analyzed text. Some institutions’ configurations differ from Turnitin’s defaults. And a “standalone” check through a different account — a friend’s instructor login, a university writing-center service, a third-party checker that licenses Turnitin’s database for similarity but not its AI model — may not even be running the analysis you think it is.

That last one deserves emphasis: plenty of “I checked it on Turnitin” stories turn out to involve a service that is not Turnitin’s AI detector at all. Comparing a Scribbr or Quetext AI score against Turnitin’s indicator is comparing two different vendors’ models, where disagreement isn’t just possible but expected.

What to actually do with a discrepancy

If you’re a student: don’t panic, and don’t chase the number. Reconstruct the inputs. Same file format both times? Same word count? Did weeks pass between checks? Was the “standalone” check genuinely Turnitin’s AI indicator or a different product? Most gaps dissolve under those four questions. Keep your drafts and version history regardless — process evidence beats score arguments every time there’s a real dispute. And if you want a pre-submission read on your prose, run your draft through a detector yourself and treat it as a forecast rather than a promise; that’s the honest use of any self-check, which is also how we frame it on our pricing page.

If you’re an instructor: a student showing you a lower self-check score isn’t necessarily gaming you, and your higher assignment-side score isn’t necessarily the truer one — they’re two noisy readings of the same object under different conditions. Turnitin’s own guidance says the indicator alone shouldn’t drive an integrity decision. Score variance across venues is one more reason that guidance exists: if the number were solid evidence, it would at least agree with itself.

The takeaway worth keeping: the score isn’t a property of your essay. It’s a property of *this text, extracted this way, scored by this model version, under these settings, on this day*. Change any clause and the number can move. There are more detection guides on the blog unpacking each of those clauses in depth.

Frequently asked questions

Is Turnitin’s AI detector different inside Canvas than on turnitin.com?

No — it’s the same engine on Turnitin’s servers either way. Canvas is just the delivery route; the LTI integration ships your file to Turnitin, which runs the identical analysis it runs on direct submissions. When scores differ between the two routes, the cause is almost always in what surrounds the analysis: the file arrived in a different format, the submission happened weeks apart across a model update, or the two assignments carry different settings. The detector itself doesn’t know or care which door your essay came through.

Why did my AI score change when I resubmitted the exact same file?

Two main reasons. First, time: Turnitin updates its detection models, so a resubmission months later is scored by a slightly different model than the original. Second, borderline text: the AI indicator is a probabilistic estimate, and passages near the decision boundary can tip either way between runs, which is why Turnitin flags lower scores with extra uncertainty. Small score movement on identical text isn’t evidence of a glitch — it’s what probabilistic classification looks like at the edges.

Does submitting a PDF instead of a Word file change the AI score?

It can. The detector analyzes extracted text, and extraction differs by format. A .docx yields clean text; a PDF — especially one exported oddly or built from images — can introduce broken line endings, merged words, dropped sections, or no text at all. Different extracted text is genuinely a different input, and a different input can score differently. If two submissions of ‘the same essay’ went in as different file types, compare the extracted text before blaming the detector.

Which score counts if Canvas and a standalone check disagree?

For your grade and any integrity process, the score attached to the actual assignment submission is the one that exists institutionally — an earlier or later check through another account is invisible to your instructor. That said, a discrepancy is worth understanding rather than fearing: note the dates, file types, and word counts involved. If a conversation ever happens, being able to say ‘the checks ran weeks apart on different file formats’ is a concrete, verifiable explanation for a gap.

Can I use a self-check score to predict my class assignment’s score?

Only loosely. A self-check through a different account or service runs at a different time, possibly on a different model version, sometimes on a different extraction of your file — so treat it as a weather forecast, not a guarantee. It will usually tell you whether you’re in clear, cloudy, or stormy territory, and that’s genuinely useful. But a 4% self-check doesn’t promise a 4% assignment score, and building your submission decisions around one exact number misreads how much noise these estimates carry.

Try it on your own text

Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.