How Does Turnitin Detect AI Writing?
Turnitin scores overlapping windows of your sentences for statistical predictability, then pools them into one number. The mechanism, from its own docs.
By the Undetected.ai team
August 2026 · 8 min read
2 free runs a day, up to 200 words each. We save your run so you can get back to it, and a delete button appears with the result. See privacy.
This is our own AI-pattern score, measured here on sentence rhythm, template phrases, vocabulary variety and passive voice. It is not a GPTZero, Turnitin, Originality.ai, Copyleaks or ZeroGPT result, and it does not predict one. Worth knowing: we also ask the rewrite to vary sentence length, drop template phrases and prefer the active voice, so some of the drop is built in. Read the two panels below, not just the number.
Before ·
After ·
Turnitin detects AI writing statistically, not by looking anything up. It breaks a submission into overlapping segments of sentences, scores each segment between 0 and 1 for how likely it is to be machine-written, gives every qualifying sentence its segment's score, pools the sentences that appear in more than one segment, then aggregates those sentence scores into the single document percentage an instructor sees. There is no database of AI text and no watermark involved.
That distinction explains almost every confusing thing about the AI writing report: why it needs 300 words before it will run, why it ignores your bullet points, why anything between 1% and 19% comes back as an asterisk, and why the number it shows is deliberately lower than what the model actually believes.
How does Turnitin detect AI writing?
Turnitin's own FAQ describes the pipeline in one paragraph, and it is worth reading closely because it is the only first-party account that exists. When a paper is submitted, sentences are extracted and segmented into overlapping sections. Each segment goes through the classifier and comes back with a value between 0 and 1, which is the probability that the text is AI-generated rather than human. Each qualifying sentence inside a segment inherits that segment's score. Because the segments overlap, some sentences end up with several scores, which are pooled into one. Those sentence scores are then aggregated into the overall document percentage.
The overlapping windows are the interesting engineering choice. A model that scored whole documents at once would be easy to fool by burying three AI paragraphs in twenty human ones. Scoring in overlapping chunks and pooling means a run of machine-written sentences raises the local score even when the rest of the paper is genuinely yours, which is how the report can highlight specific passages rather than just returning one number.
What is Turnitin's model actually looking for?
Word probability. Turnitin's explanation is that large language models generate text by repeatedly choosing highly probable next words, so their output tends to be consistent and predictable in a way human writing is not. Human writing, in its wording, "tends to be inconsistent and idiosyncratic, resulting in a low probability of picking the next word the human will use in the sequence." The classifiers are trained to separate those two signatures.
This is the same underlying idea behind every statistical AI detector, and it has a consequence people rarely follow through on. The detector is not measuring whether a machine was involved. It is measuring how predictable your prose is. Writing that is careful, formulaic, heavily edited toward clarity, or produced by someone writing in a second language can all land in the predictable region without any AI having touched it. Turnitin says it trained against that bias deliberately, sampling second-language learners, students at institutions with diverse enrollments, and less common subject areas such as anthropology, geology and sociology.
What counts as qualifying text?
Only prose sentences inside paragraphs. Turnitin calls this qualifying text and defines it as individual sentences contained in paragraphs that make up a longer piece of written work, such as an essay, a dissertation or an article. Everything else is excluded from the analysis: poetry, scripts, code, bullet points, tables and annotated bibliographies.
That exclusion is the reason the percentage and the highlights sometimes seem to disagree. If half your document is a table and the model only scored the prose, the percentage describes the prose, not the page. Turnitin says so directly: a document containing several different writing types will produce a disparity between the percentage and the highlights.
There are hard file limits too, and they are stricter than most people expect:
| Requirement | Turnitin's rule |
|---|---|
| Minimum length | At least 300 words of prose in a long-form format |
| Maximum length | No more than 30,000 words |
| Languages | English, Spanish and Japanese only |
| File types | .docx, .pdf, .txt, .rtf |
| File size | Under 100 MB |
Miss any of these and no report is generated at all. A submission in an unsupported language returns an empty error state rather than a score. This also catches out anyone submitting a scan rather than a document: a photographed or scanned page has no text layer for Turnitin to read, so you would need to pull the text out of the scanned PDF before any analysis of any kind could run on it.
Why does Turnitin hide scores between 1% and 19%?
Because its own testing showed that band is unreliable. Turnitin states that "no score or highlights are attributed for AI detection scores in the 1% to 19% range", and gives the reason plainly: its testing "found that there is a higher incidence of false positives when the percentage is between 0 and 19". Instead of a number, the report displays an asterisk.
This is a genuinely unusual thing for a vendor to do. Turnitin is choosing to publish less information about its own product in order to reduce the chance of someone being wrongly accused. Reports generated before July 8, 2024 may still show a numeric score below 20%, which is worth knowing if you are looking at an old report. We have covered the practical side of that threshold in our guide to what percentage of AI detection is allowed in Turnitin.
Does Turnitin detect paraphrasing tools and AI humanizers?
Turnitin says yes, and names one. Its documentation states the model detects qualifying text likely generated by an LLM and also detects "when likely AI generated text may have been further modified by an AI bypasser, AI-paraphrasing tool or AI word spinner, such as Quillbot". That text is highlighted in the report the same way plain AI output is.
Two caveats sit alongside that claim. The first is that it applies to English only; Turnitin states its Spanish and Japanese detectors do not include AI paraphrasing or bypasser detection. The second is that this is a vendor claim about its own capability, with no independent test behind it, exactly like the accuracy figures. We sell a rewriting tool and we will still say the obvious thing: no humanizer can promise you a Turnitin result, because no humanizer can see your institution's Turnitin instance. The full picture is in our breakdown of how accurate the Turnitin AI detector really is.
Why is the reported percentage lower than the model's real estimate?
Because Turnitin tuned it that way on purpose, and published the size of the effect. To hold false positives under 1%, it accepts missing some AI text, and it gives a worked example: "if we identify that 50% of a document is likely written by an AI tool, it could contain as much as 65% AI writing."
So the percentage is a floor rather than a best guess. For an instructor that means a low number is weak evidence of anything, and a high number is more serious than it looks. For a student it cuts the other way: a clean report is not proof of a clean paper, it is proof that the model did not clear its own confidence bar.
How does Turnitin's AI score relate to the similarity score?
It does not. Turnitin is explicit that the AI writing percentage is "different from and independent of the similarity score", and that AI writing highlights are not visible in the Similarity Report at all. They measure unrelated things: similarity finds text matching sources that already exist, AI detection estimates whether a machine composed text that is otherwise original.
A paper can score 0% similarity and a high AI percentage, which is the normal outcome for something written by a model, since machine-composed sentences are not copied from anywhere. It can equally score high similarity and no AI flag, which is the normal outcome for a human writer quoting heavily.
This trips people up most often inside a learning management system, because the number the platform surfaces is the similarity figure. Turnitin's Canvas documentation says the Gradebook shows an indication of the Similarity Report percentage, while the AI indicator stays inside the report itself and is visible to instructors only. So a percentage you can see is a matching score by definition, a point we unpack alongside the rest of the setup on our Canvas AI detector page.
Two further Turnitin features are also separate from AI detection and often confused with it. Authorship, formerly Originality, uses document metadata and forensic language analysis to assess whether a submission was written by someone other than the student, and Turnitin says it cannot indicate whether text was AI written. Turnitin Clarity's Writing Report is different again: it records pasting activity, replays a timeline of the document being written, and surfaces the student's AI chat interactions when the assistant is enabled. That is provenance rather than statistics, and no amount of rewriting affects it.
Which AI models can Turnitin detect?
Turnitin publishes the list, and it is current. For English submissions it names output from GPT-5.4 and GPT-5.4-pro, GPT-5.3, the GPT-5 family, GPT-4o, Claude Sonnet-4.6, Claude Opus-4.5, Claude Haiku-4.5, Gemini-3.1-pro, Gemini-3-pro, Grok-4.1, LLaMA-4-Maverick, Mistral-Large-3, Deepseek-v3.2 and about a dozen others, plus tools built on top of them.
The practical reading is that no frontier model is currently outside the list, so advice built on picking a less-detected model is weaker than it sounds. Turnitin also changed its architecture in July 2026, consolidating what had been a multi-model ensemble into a single model while stating it maintains the same sub-1% false positive rate. Any article about Turnitin's detection method written before that date is describing a system that no longer exists.
The short version
Turnitin scores overlapping windows of your sentences for statistical predictability, pools those scores, and reports one number that is deliberately conservative. It only reads prose, only in three languages, only between 300 and 30,000 words, and only shows the instructor a figure once it passes 20%. It claims to catch paraphrasers and bypassers in English. It does not compare your work to a library of AI text, because no such library exists, and it does not know whether you used a model. It has an opinion about how predictable your writing is, and that opinion is the entire mechanism.
Let Undetected.ai clear the flag for you
Paste your own text and watch our AI-pattern gauge sweep from the score on your draft to the score on the rewrite, meaning kept intact.