Undetected.ai

Updated August 2026 · every figure traced to its source

AI detector false positive rate: how often AI detection wrongly flags human writing

Published false positive rates for the same handful of detectors range from 0.004 percent to 61.3 percent. That is a spread of four orders of magnitude for tools doing the same job.

The detectors are not the variable. Below is every claim set against what outside testing found, and the three things that actually decide which number you get.

The Humanizer

Try:
Tone
Strength:

2 free runs a day, up to 200 words each. We save your run so you can get back to it, and a delete button appears with the result. See privacy.

AI-pattern score

This is our own AI-pattern score, measured here on sentence rhythm, template phrases, vocabulary variety and passive voice. It is not a GPTZero, Turnitin, Originality.ai, Copyleaks or ZeroGPT result, and it does not predict one. Worth knowing: we also ask the rewrite to vary sentence length, drop template phrases and prefer the active voice, so some of the drop is built in. Read the two panels below, not just the number.

Before ·

After ·

·

Your own text rewritten · Saved to your history, delete any time

The short answer

There is no single AI detector false positive rate, and the range is not a rounding difference. Peer-reviewed testing on US undergraduate essays measured about 1.3 percent, which was lower than the 5.0 percent error rate of the human graders reading the same work. Testing on TOEFL essays by non-native English speakers measured 61.3 percent with the same class of tools. Vendors publish figures between 0.004 percent and 1 percent, generated on data they selected. Three variables explain almost the entire spread: how long the passage is, who wrote it, and who ran the test. Short text and non-native writing produce the high numbers; long, native-English, clean prose produces the low ones. Any institution acting on a single score, on a short submission, without asking which detector and at what threshold, is working with a number that does not support the decision being made.

Last updated August 4, 2026. Vendor claims were read directly from vendor documentation on that date. Independent figures are peer-reviewed or attributed to a named study. One widely circulated figure is excluded, and we explain why below.

Claim versus measurement

What each detector claims, and what outside testing found

Vendor figures and independent figures are kept in separate columns on purpose. They are not the same kind of number and merging them is how this topic became so confused.

Detector Vendor's own claim Where that claim comes from Independently measured Source of that measurement
Turnitin Under 1% (marketing), 0.51% at document level (whitepaper) Turnitin's own blog and whitepaper About 50% in a small-sample press test Washington Post, April 1 2023
Copyleaks 0.03% on English (99.97% of human text classed human) Copyleaks per-language table, read Aug 4 2026 No independent measurement published None found
GPTZero 0.9% (99.1% of human articles classed human at high confidence) GPTZero's own FAQ At or below 1% on medium and longer passages Chicago Booth working paper, Aug 2025
Originality.ai Not published as a single false positive figure Not stated At or below 1% on medium and longer passages Chicago Booth working paper, Aug 2025
Pangram 0.004% on academic essays, about 1 in 10,000 overall Pangram's own testing Essentially zero on medium and longer passages Chicago Booth working paper, Aug 2025

Turnitin

Also refuses to score under 300 words and suppresses any reading between 1 and 19 percent

Copyleaks

Markets a 0.2% figure elsewhere, so even its own two numbers differ by nearly 7x

GPTZero

The claim holds only at high confidence, which is a subset of all verdicts

Originality.ai

Its own guidance says to use detection as one signal, not a final decision

Pangram

The lowest published rate in the category, and the only vendor claim an outside test broadly supported

Turnitin is the row worth staring at. Its own two published numbers differ, and the one independent test on record differs from both by a factor of roughly 100.

Three variables, not one

Why the published rates disagree by four orders of magnitude

The detectors are broadly similar technology. Nearly all of the variance comes from what they were pointed at.

Variable 1

Length of the passage

This is the single biggest variable and the least discussed. Detectors work from statistical regularity, and a short passage simply does not contain enough of it, so verdicts swing hard. The Chicago Booth work found rates at or below 1 percent on medium and long passages for several tools, and materially worse behaviour on short ones. Turnitin sidesteps the problem by refusing to score anything under 300 words at all.

Variable 2

Who wrote the text

The gap between 1.3 percent on US undergraduate STEM essays and 61.3 percent on TOEFL essays is the same technology measured on two populations. Non-native English writers, and anyone trained to write in plain, even, uniform prose, sit closer to the machine profile. Academic and technical conventions push writing in exactly that direction.

Variable 3

Who ran the test

Vendor numbers are generated on data the vendor chose, at a confidence threshold the vendor chose, and are almost never reproducible. That does not make them lies, but it does make them best cases. Where an outside test exists, it is usually less flattering than the vendor figure, and the two are rarely comparable because the inputs differ.

Put those together and the headline numbers stop being contradictory. A vendor testing long, clean, native-English academic prose and reporting 0.5 percent, and a research team testing TOEFL essays and reporting 61.3 percent, can both be accurate. They measured different things and only one of them resembles the full range of writing a real institution processes.

The published evidence

What the peer-reviewed research found

Three studies carry most of the weight on this question. None of them is quoted in the vendor comparisons that rank for it.

Advances in Physiology Education, 2025

Detectors made fewer mistakes than the professors

Hyatt and colleagues had 190 undergraduates in an anatomy and physiology course hand-write essays, then ran 50 of those essays through four AI detectors and 48 of them past nine human raters. The detectors produced false positives on about 1.3 percent of the essays. The human raters produced false positives on 5.0 percent. The three best detectors identified work correctly 93 to 98 percent of the time against 84 to 95 percent for the humans. Running the detectors in aggregate, and treating only a consensus as meaningful, cut the false positive likelihood to nearly zero.

This is the strongest evidence that a single detector reading is the problem rather than detection itself. It is also uncomfortable for both sides of the argument: the tools were wrong, and the people checking by eye were wrong more often.

Hyatt et al., Adv Physiol Educ 2025 (PubMed 40105702)

Patterns, 2023

The same detectors were wrong 61.3% of the time on one population

Liang and colleagues at Stanford ran seven detectors over TOEFL essays written by non-native English speakers. On average 61.3 percent of those genuinely human essays were classified as AI-generated, and every one of the seven detectors misclassified more than half of them. The same detectors were close to perfect on essays by native English speakers.

Nothing about the detectors changed between those two results. Only the writers did. Simpler, more regular sentence construction reads as low perplexity, which is the exact signal being scored.

Liang et al., Patterns 2023

ACL 2024

At a strict false positive budget, most detectors stop working

RAID is the largest benchmark in this field: over 6 million generations across 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies, used to evaluate 12 detectors. Held to a false positive rate below 1 percent, which is roughly the threshold an institution would need before accusing anyone, most of the detectors tested could not operate. The authors also found detectors are "easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models."

A detector can be accurate on average and still be unusable for decisions, because the decision threshold that protects innocent writers is stricter than the one that maximizes accuracy.

RAID benchmark, arXiv:2405.07940

Read this before you trust a percentage

The figure we will not repeat

If you search this topic you will quickly meet a precise, confident number: ZeroGPT has a 24.64 percent false positive rate. It appears on comparison pages, in listicles, and in articles that cite each other.

We tried to trace it. We could not find a published study, a sample size, a date, a text corpus or a method behind it anywhere. So it is not in the table above, and it will not be on this page. That is not a claim that ZeroGPT performs well, and we have written separately about ZeroGPT's accuracy. It is a statement that a number with two decimal places and no method is not evidence, and repeating it would make this page worse.

This pattern is the norm in this niche rather than the exception. The per-model detection rates people quote, the Grammarly false positive figures, and this one all share a shape: a suspiciously exact percentage, no methodology, and a trail of citations leading back to companies selling detection or evasion software. Both sides of this market have a commercial interest in a specific answer.

The rule we hold ourselves to: a figure goes on this site only if it comes from a peer-reviewed paper, a named independent test, or a vendor's own documentation read directly and labelled as a vendor claim. Where a number cannot be traced, we say that instead of repeating it. We sell a humanizer, so treat our reasoning with the same suspicion, and check the sources linked throughout.

Practical consequences

What a false positive rate actually means for you

If you are being assessed

A 1 percent false positive rate sounds tolerable until you apply it at scale. Across 2,000 submissions in a term it means roughly 20 innocent people flagged. Ask which detector was used, on how many words, and at what confidence. A verdict on a 200 word passage is not a result worth defending against.

If you are doing the assessing

The Hyatt study points at the fix. One detector reading is unreliable, agreement between several is far stronger, and both Originality.ai and Grammarly state in their own documentation that detection should not be used as a standalone verification method. Provenance beats probability: version history and authorship records show process rather than guessing at it.

If you write in the flagged style

Plain, uniform, carefully edited prose is the profile that gets caught, which penalizes exactly the writers who were trained well. Varying sentence length and adding concrete specifics moves the measurement, because those are the properties being measured. That is what our AI humanizer changes.

One honest limit. Rewriting changes what a detector measures, because a detector only ever sees finished text. It does not touch provenance features that record how a document was written, such as Grammarly Authorship or GPTZero Writing Reports, which replay composition keystroke by keystroke. If you are answering a false accusation, that record is your evidence and a rewrite is not. We would rather say so than sell you the wrong remedy. There is more in our guide to proving you did not use AI.

Questions people ask

AI detector false positives, answered

What is the false positive rate of AI detectors?

There is no single rate, and any page quoting one number is oversimplifying. Published figures for the same detectors run from 0.004 percent to 61.3 percent depending on how long the text is and who wrote it. For long, native-English writing the better tools measure at or below 1 percent in independent testing. For short passages, and for essays by non-native English speakers, the rate rises dramatically.

How often do AI detectors give false positives?

On typical US student coursework of reasonable length, independent peer-reviewed testing puts it near 1.3 percent, which was lower than the 5.0 percent error rate of human graders reading the same essays. That figure climbs steeply for short submissions and for writers who learned English as a second language, where one study measured 61.3 percent.

What is Turnitin's false positive rate?

Turnitin states under 1 percent in its marketing and 0.51 percent at document level in its whitepaper. A 2023 Washington Post test on a small sample produced roughly 50 percent. Both figures are real and they are not comparable, because they used different text. Turnitin also refuses to score submissions under 300 words and suppresses any reading between 1 and 19 percent, which hides the band where its false positives cluster.

Which AI detector has the lowest false positive rate?

On the evidence available, Pangram. It claims about 0.004 percent on academic essays, and it is the only vendor whose low claim was broadly supported by an outside test, the 2025 Chicago Booth working paper, which found essentially zero false positives on medium and longer passages. GPTZero and Originality.ai measured at or below 1 percent on the same passages in that work.

Can AI detectors be wrong?

Yes, in both directions, and every major vendor says so on its own site. They miss AI text and they flag human text. Turnitin has been documented missing roughly 15 percent of AI-generated content within a document. The important point for anyone being judged by one is that a detector reports a statistical probability about the text, not evidence about who wrote it.

Why do AI detectors flag human writing as AI?

Because they measure predictability rather than authorship. Two signals do most of the work: perplexity, meaning how expected each word is, and burstiness, meaning how much sentence length varies. Clear, measured, uniform prose scores low on both. That describes machine output, and it also describes a well-drilled academic or technical writer, which is why competent human writing gets caught.

Do AI detectors discriminate against non-native English speakers?

The measured effect is large. Stanford researchers found seven detectors flagged 61.3 percent of TOEFL essays by non-native English writers as AI-generated, while classifying native-speaker essays almost perfectly. The cause is mechanical rather than deliberate: a smaller working vocabulary and more regular sentence construction produce exactly the low-perplexity pattern detectors score as machine-like.

What should I do if an AI detector falsely flags my writing?

Lead with provenance rather than argument. Version history in Google Docs or Word, and tools that record how a document was composed such as Grammarly Authorship or GPTZero Writing Reports, show the writing process itself, which a detector score cannot rebut. Ask which detector was used, on how many words, and at what threshold. A score on a short passage is close to meaningless.

Write so the question never comes up

Detectors score predictability and rhythm. If your writing is clear, even and uniform, it scores like a machine no matter who typed it. Our humanizer changes the properties being measured while keeping your meaning, your facts and your argument intact.

Plans planned from $9/mo · Free account to start · Runs saved to your history, delete any time

Humanize my text