Updated August 2026 · every figure traced to its source
AI detector false positive rate: how often AI detection wrongly flags human writing
Published false positive rates for the same handful of detectors range from 0.004 percent to 61.3 percent. That is a spread of four orders of magnitude for tools doing the same job.
The detectors are not the variable. Below is every claim set against what outside testing found, and the three things that actually decide which number you get.
2 free runs a day, up to 200 words each. We save your run so you can get back to it, and a delete button appears with the result. See privacy.
This is our own AI-pattern score, measured here on sentence rhythm, template phrases, vocabulary variety and passive voice. It is not a GPTZero, Turnitin, Originality.ai, Copyleaks or ZeroGPT result, and it does not predict one. Worth knowing: we also ask the rewrite to vary sentence length, drop template phrases and prefer the active voice, so some of the drop is built in. Read the two panels below, not just the number.
Before ·
After ·
Your own text rewritten · Saved to your history, delete any time
The short answer
There is no single AI detector false positive rate, and the range is not a rounding difference. Peer-reviewed testing on US undergraduate essays measured about 1.3 percent, which was lower than the 5.0 percent error rate of the human graders reading the same work. Testing on TOEFL essays by non-native English speakers measured 61.3 percent with the same class of tools. Vendors publish figures between 0.004 percent and 1 percent, generated on data they selected. Three variables explain almost the entire spread: how long the passage is, who wrote it, and who ran the test. Short text and non-native writing produce the high numbers; long, native-English, clean prose produces the low ones. Any institution acting on a single score, on a short submission, without asking which detector and at what threshold, is working with a number that does not support the decision being made.
Last updated August 4, 2026. Vendor claims were read directly from vendor documentation on that date. Independent figures are peer-reviewed or attributed to a named study. One widely circulated figure is excluded, and we explain why below.
Claim versus measurement
What each detector claims, and what outside testing found
Vendor figures and independent figures are kept in separate columns on purpose. They are not the same kind of number and merging them is how this topic became so confused.
| Detector | Vendor's own claim | Where that claim comes from | Independently measured | Source of that measurement |
|---|---|---|---|---|
| Turnitin | Under 1% (marketing), 0.51% at document level (whitepaper) | Turnitin's own blog and whitepaper | About 50% in a small-sample press test | Washington Post, April 1 2023 |
| Copyleaks | 0.03% on English (99.97% of human text classed human) | Copyleaks per-language table, read Aug 4 2026 | No independent measurement published | None found |
| GPTZero | 0.9% (99.1% of human articles classed human at high confidence) | GPTZero's own FAQ | At or below 1% on medium and longer passages | Chicago Booth working paper, Aug 2025 |
| Originality.ai | Not published as a single false positive figure | Not stated | At or below 1% on medium and longer passages | Chicago Booth working paper, Aug 2025 |
| Pangram | 0.004% on academic essays, about 1 in 10,000 overall | Pangram's own testing | Essentially zero on medium and longer passages | Chicago Booth working paper, Aug 2025 |
Turnitin
Also refuses to score under 300 words and suppresses any reading between 1 and 19 percent
Copyleaks
Markets a 0.2% figure elsewhere, so even its own two numbers differ by nearly 7x
GPTZero
The claim holds only at high confidence, which is a subset of all verdicts
Originality.ai
Its own guidance says to use detection as one signal, not a final decision
Pangram
The lowest published rate in the category, and the only vendor claim an outside test broadly supported
Turnitin is the row worth staring at. Its own two published numbers differ, and the one independent test on record differs from both by a factor of roughly 100.
Three variables, not one
Why the published rates disagree by four orders of magnitude
The detectors are broadly similar technology. Nearly all of the variance comes from what they were pointed at.
Variable 1
Length of the passage
This is the single biggest variable and the least discussed. Detectors work from statistical regularity, and a short passage simply does not contain enough of it, so verdicts swing hard. The Chicago Booth work found rates at or below 1 percent on medium and long passages for several tools, and materially worse behaviour on short ones. Turnitin sidesteps the problem by refusing to score anything under 300 words at all.
Variable 2
Who wrote the text
The gap between 1.3 percent on US undergraduate STEM essays and 61.3 percent on TOEFL essays is the same technology measured on two populations. Non-native English writers, and anyone trained to write in plain, even, uniform prose, sit closer to the machine profile. Academic and technical conventions push writing in exactly that direction.
Variable 3
Who ran the test
Vendor numbers are generated on data the vendor chose, at a confidence threshold the vendor chose, and are almost never reproducible. That does not make them lies, but it does make them best cases. Where an outside test exists, it is usually less flattering than the vendor figure, and the two are rarely comparable because the inputs differ.
Put those together and the headline numbers stop being contradictory. A vendor testing long, clean, native-English academic prose and reporting 0.5 percent, and a research team testing TOEFL essays and reporting 61.3 percent, can both be accurate. They measured different things and only one of them resembles the full range of writing a real institution processes.
The published evidence
What the peer-reviewed research found
Three studies carry most of the weight on this question. None of them is quoted in the vendor comparisons that rank for it.
Advances in Physiology Education, 2025
Detectors made fewer mistakes than the professors
Hyatt and colleagues had 190 undergraduates in an anatomy and physiology course hand-write essays, then ran 50 of those essays through four AI detectors and 48 of them past nine human raters. The detectors produced false positives on about 1.3 percent of the essays. The human raters produced false positives on 5.0 percent. The three best detectors identified work correctly 93 to 98 percent of the time against 84 to 95 percent for the humans. Running the detectors in aggregate, and treating only a consensus as meaningful, cut the false positive likelihood to nearly zero.
This is the strongest evidence that a single detector reading is the problem rather than detection itself. It is also uncomfortable for both sides of the argument: the tools were wrong, and the people checking by eye were wrong more often.
Patterns, 2023
The same detectors were wrong 61.3% of the time on one population
Liang and colleagues at Stanford ran seven detectors over TOEFL essays written by non-native English speakers. On average 61.3 percent of those genuinely human essays were classified as AI-generated, and every one of the seven detectors misclassified more than half of them. The same detectors were close to perfect on essays by native English speakers.
Nothing about the detectors changed between those two results. Only the writers did. Simpler, more regular sentence construction reads as low perplexity, which is the exact signal being scored.
ACL 2024
At a strict false positive budget, most detectors stop working
RAID is the largest benchmark in this field: over 6 million generations across 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies, used to evaluate 12 detectors. Held to a false positive rate below 1 percent, which is roughly the threshold an institution would need before accusing anyone, most of the detectors tested could not operate. The authors also found detectors are "easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models."
A detector can be accurate on average and still be unusable for decisions, because the decision threshold that protects innocent writers is stricter than the one that maximizes accuracy.
Read this before you trust a percentage
The figure we will not repeat
If you search this topic you will quickly meet a precise, confident number: ZeroGPT has a 24.64 percent false positive rate. It appears on comparison pages, in listicles, and in articles that cite each other.
We tried to trace it. We could not find a published study, a sample size, a date, a text corpus or a method behind it anywhere. So it is not in the table above, and it will not be on this page. That is not a claim that ZeroGPT performs well, and we have written separately about ZeroGPT's accuracy. It is a statement that a number with two decimal places and no method is not evidence, and repeating it would make this page worse.
This pattern is the norm in this niche rather than the exception. The per-model detection rates people quote, the Grammarly false positive figures, and this one all share a shape: a suspiciously exact percentage, no methodology, and a trail of citations leading back to companies selling detection or evasion software. Both sides of this market have a commercial interest in a specific answer.
The rule we hold ourselves to: a figure goes on this site only if it comes from a peer-reviewed paper, a named independent test, or a vendor's own documentation read directly and labelled as a vendor claim. Where a number cannot be traced, we say that instead of repeating it. We sell a humanizer, so treat our reasoning with the same suspicion, and check the sources linked throughout.
Practical consequences
What a false positive rate actually means for you
If you are being assessed
A 1 percent false positive rate sounds tolerable until you apply it at scale. Across 2,000 submissions in a term it means roughly 20 innocent people flagged. Ask which detector was used, on how many words, and at what confidence. A verdict on a 200 word passage is not a result worth defending against.
If you are doing the assessing
The Hyatt study points at the fix. One detector reading is unreliable, agreement between several is far stronger, and both Originality.ai and Grammarly state in their own documentation that detection should not be used as a standalone verification method. Provenance beats probability: version history and authorship records show process rather than guessing at it.
If you write in the flagged style
Plain, uniform, carefully edited prose is the profile that gets caught, which penalizes exactly the writers who were trained well. Varying sentence length and adding concrete specifics moves the measurement, because those are the properties being measured. That is what our AI humanizer changes.
One honest limit. Rewriting changes what a detector measures, because a detector only ever sees finished text. It does not touch provenance features that record how a document was written, such as Grammarly Authorship or GPTZero Writing Reports, which replay composition keystroke by keystroke. If you are answering a false accusation, that record is your evidence and a rewrite is not. We would rather say so than sell you the wrong remedy. There is more in our guide to proving you did not use AI.
Questions people ask
AI detector false positives, answered
What is the false positive rate of AI detectors?
There is no single rate, and any page quoting one number is oversimplifying. Published figures for the same detectors run from 0.004 percent to 61.3 percent depending on how long the text is and who wrote it. For long, native-English writing the better tools measure at or below 1 percent in independent testing. For short passages, and for essays by non-native English speakers, the rate rises dramatically.
How often do AI detectors give false positives?
On typical US student coursework of reasonable length, independent peer-reviewed testing puts it near 1.3 percent, which was lower than the 5.0 percent error rate of human graders reading the same essays. That figure climbs steeply for short submissions and for writers who learned English as a second language, where one study measured 61.3 percent.
What is Turnitin's false positive rate?
Turnitin states under 1 percent in its marketing and 0.51 percent at document level in its whitepaper. A 2023 Washington Post test on a small sample produced roughly 50 percent. Both figures are real and they are not comparable, because they used different text. Turnitin also refuses to score submissions under 300 words and suppresses any reading between 1 and 19 percent, which hides the band where its false positives cluster.
Which AI detector has the lowest false positive rate?
On the evidence available, Pangram. It claims about 0.004 percent on academic essays, and it is the only vendor whose low claim was broadly supported by an outside test, the 2025 Chicago Booth working paper, which found essentially zero false positives on medium and longer passages. GPTZero and Originality.ai measured at or below 1 percent on the same passages in that work.
Can AI detectors be wrong?
Yes, in both directions, and every major vendor says so on its own site. They miss AI text and they flag human text. Turnitin has been documented missing roughly 15 percent of AI-generated content within a document. The important point for anyone being judged by one is that a detector reports a statistical probability about the text, not evidence about who wrote it.
Why do AI detectors flag human writing as AI?
Because they measure predictability rather than authorship. Two signals do most of the work: perplexity, meaning how expected each word is, and burstiness, meaning how much sentence length varies. Clear, measured, uniform prose scores low on both. That describes machine output, and it also describes a well-drilled academic or technical writer, which is why competent human writing gets caught.
Do AI detectors discriminate against non-native English speakers?
The measured effect is large. Stanford researchers found seven detectors flagged 61.3 percent of TOEFL essays by non-native English writers as AI-generated, while classifying native-speaker essays almost perfectly. The cause is mechanical rather than deliberate: a smaller working vocabulary and more regular sentence construction produce exactly the low-perplexity pattern detectors score as machine-like.
What should I do if an AI detector falsely flags my writing?
Lead with provenance rather than argument. Version history in Google Docs or Word, and tools that record how a document was composed such as Grammarly Authorship or GPTZero Writing Reports, show the writing process itself, which a detector score cannot rebut. Ask which detector was used, on how many words, and at what threshold. A score on a short passage is close to meaningless.
Keep reading
Related guides
Why AI detectors flag human writing
The mechanism behind a false positive, who gets caught most, and how to fix a flag the right way.
Which AI detector is most accurate
How the five major detectors compare on accuracy, and what their published numbers leave out.
Best AI humanizer for Turnitin
What Turnitin actually reports, the 300 word floor, and the suppressed 1 to 19 percent band.
Best AI humanizer for Copyleaks
Copyleaks publishes the lowest false positive claim of the mainstream detectors. What that figure covers.
Undetectable AI pricing compared
Verified current pricing across the major humanizers, including the per-input word caps most of them do not advertise.
Best AI humanizer
The full comparison of humanizers, what each one costs, and which detectors they are tuned against.
Does Google detect AI content?
The same tracing method applied to search: what Google has actually published, and the LinkedIn line the industry treats as proof.
Write so the question never comes up
Detectors score predictability and rhythm. If your writing is clear, even and uniform, it scores like a machine no matter who typed it. Our humanizer changes the properties being measured while keeping your meaning, your facts and your argument intact.
Plans planned from $9/mo · Free account to start · Runs saved to your history, delete any time