Turnitin AI Score: What the Percentage Actually Means

The number measures text, not you. Turnitin will not say what counts as too high, and it hides everything below 20% because that range produced too many false positives.

By HumanizeBot Editorial Team Published 7 min read Editorial policy
Turnitin AI Score: What the Percentage Actually Means

You open your submission receipt and there it is: a percentage next to the words "AI writing." Maybe it says 34%. Maybe it says 81%. Either way, your stomach drops, and the first question is always the same — how much trouble am I in?

Before anything else, understand what that number is and is not. It is an estimate produced by a statistical model about a block of text. It is not a verdict, it is not evidence, and it is not a measurement of you. Turnitin itself is unusually direct about this: the company states that the AI writing indicator is not proof of misconduct and should not be used on its own to accuse a student. Knowing exactly what the score measures is the difference between handling a flag calmly and making the situation worse.

The single most common confusion: AI score is not similarity score

Turnitin produces two completely separate numbers, and students mix them up constantly.

The similarity score is the older, familiar one. It compares your text against a giant index of web pages, publications, and previously submitted student papers, and reports what proportion of your document matches something else. A 22% similarity score usually means quoted material, a bibliography, and common phrasing — not cheating.

The AI writing indicator works on a totally different principle. It does not compare your text to anything. It has no database of "AI-written documents" to match against. Instead, it analyses the statistical texture of your prose — word choice patterns, sentence construction, predictability — and estimates how much of it resembles the output of a large language model.

 Similarity scoreAI writing indicator
What it doesMatches your text against an index of existing sourcesStatistically models how "machine-like" your prose reads
Can you see why?Yes — it shows you the matching sourceNo — it highlights passages but cannot show a source
Verifiable?Yes, you can open the matched documentNo, there is nothing external to check
Means misconduct?Not automatically — depends on citationNot automatically — Turnitin says so explicitly

That third row matters enormously. A similarity flag is falsifiable: you can open the source it matched and see whether you quoted it properly. An AI flag is not falsifiable in the same way. There is no document to point at. This is precisely why it carries far less evidentiary weight, and why a careful instructor will treat it as a prompt to have a conversation rather than a finding of fact.

What the percentage is actually counting

The number is not a confidence level. It is not "we are 60% sure this is AI." It is an estimate of the proportion of qualifying text in your document that the model believes was AI-generated.

The phrase "qualifying text" is doing a lot of work there. Turnitin's AI indicator only assesses long-form prose written in English. Several things in a typical assignment are excluded from the calculation entirely:

  • Bullet points and short list items
  • Code blocks
  • Tables and figure captions
  • Reference lists and bibliographies
  • Very short documents — a piece needs a meaningful volume of continuous prose before the model will produce a score at all
  • Text in languages other than English

This produces a counterintuitive result worth knowing. If your 2,000-word report contains 1,200 words of tables, lists, and references, the percentage is calculated against the remaining 800 words of prose, not the full document. A "50% AI" flag on that paper refers to roughly 400 words — not 1,000.

Why scores under 20% are hidden

You will notice Turnitin does not display low scores. Below a certain band, the indicator shows nothing rather than a small number. The reason is straightforward and, to Turnitin's credit, openly stated: the low end of the range produced too many false positives to be useful. Human-written text frequently scores a few percent for no meaningful reason. Rather than have instructors chase noise, the tool suppresses it.

Think about what that admission implies. The vendor knows the model produces false positives, has quantified where they cluster, and has built a guard against the worst of them. That is responsible engineering — but it is also a clear signal that the scores above the cutoff are not immune to the same failure mode. They are just less likely to be wrong, not incapable of it.

The asterisk

Sometimes a score appears with an asterisk beside it. This marks lower confidence — typically on shorter documents where there is less prose for the model to work with. If your score carries an asterisk, the tool is telling you and your instructor that it is operating near the edge of its reliable range.

There is no official "too high"

Students ask constantly: what score gets you in trouble? Turnitin does not publish a threshold, and this is deliberate rather than evasive. The company's position is that the score is one data point for an instructor to consider alongside everything else they know — your previous work, your participation, your ability to discuss what you wrote.

In practice, thresholds get invented locally. A department decides 20% triggers a conversation. An instructor decides anything above 50% goes to the academic integrity office. These are institutional policies, not vendor guidance, and they vary wildly between schools and even between courses in the same school. If you want to know what number matters where you study, the answer is in your institution's academic integrity policy, not in Turnitin's documentation.

A percentage is a starting point for a conversation about your process. It is not the conclusion of one.

If the score on your own work is wrong

False positives happen. They happen most often to writers whose prose is clean, structured, and consistent — which describes a lot of careful students, and disproportionately describes people writing in a second language, who tend to use more common vocabulary and more regular sentence patterns. A polished, well-organised essay written entirely by hand can read as low-variance to a model that treats low variance as a machine signal.

Evidence you can actually produce

The strongest defence is process evidence, and the good news is that most writers already generate it without meaning to. Gather what you have:

  • Version history. Google Docs keeps a full revision timeline under File → Version history. Microsoft Word does the same for files stored in OneDrive or SharePoint. This shows the document growing over hours or days, with false starts and deletions — something generated text does not have.
  • Earlier drafts. Separate files, emailed copies, anything with a different timestamp.
  • Notes and outlines. Handwritten notes photographed on your phone, a mind map, a reading list with annotations.
  • Sources you consulted. Library loan records, browser history, PDFs in a folder, highlighted readings.
  • Communication. Emails to your instructor asking clarifying questions, messages to a study group, a writing centre appointment.

None of these individually prove authorship. Together, they describe a process, and a process is very hard to fake retroactively.

How to raise it with an instructor

Tone matters more than argument here. Instructors who flag AI scores are usually not out to get anyone — they have a number on a screen and no idea what to do with it. Give them a way out that is not a confrontation.

  1. Ask for a meeting rather than sending a long defensive email. Written arguments escalate; conversations de-escalate.
  2. Ask what the concern is before defending anything. Sometimes the answer is "I just wanted to check in," and a five-minute chat ends it.
  3. Ask what evidence exists beyond the score. Politely, genuinely. If the answer is "just the score," that is significant, because the vendor says the score alone is not sufficient.
  4. Offer your process evidence unprompted. "I still have the version history and my notes — would it help if I brought those?"
  5. Offer to discuss the content. Nothing settles an authorship question faster than being able to explain your argument, defend your source choices, and say why you cut a paragraph.
  6. Know your institution's process. Most have a formal appeals route and often a student advocate or ombudsman. Ask early what it is, even if you hope not to use it.

Why panicking makes it worse

The instinct after a flag is to rewrite everything immediately. Resist it, for two practical reasons.

First, if you rewrite the submitted file, you damage the very evidence that would clear you. The version history that showed steady, human drafting now ends with a frantic full-document overwrite hours after the flag. That looks worse than the original score.

Second, random rewriting does not reliably move the number. Swapping words for synonyms and chopping sentences at random tends to produce prose that is worse to read and no less machine-like statistically. People often make things worse. If you genuinely want to understand the mechanics of why detectors respond to some edits and not others, we have written about what bypassing AI detection actually involves — the honest version, including where it does not work.

The one thing worth doing immediately is preserving evidence: download your version history, save copies of drafts and notes somewhere safe, and take screenshots of anything that might change. Do that first, before you touch a word of the text.

The short version

The AI writing indicator estimates how much of your qualifying English prose resembles model output. It is separate from similarity. It is calculated on a subset of your document. It hides low scores because they were unreliable. It carries no official threshold, and its own vendor says it is not proof of anything. An asterisk means the tool is less sure than usual.

If you wrote your work, you have a process behind it, and that process is your strongest answer. Preserve it, stay calm, and treat the flag as a question to answer rather than a sentence to serve.

More from the blog