Undetected.ai
All posts
Detectors

Copyleaks False Positive Rate: How Accurate Is It?

Copyleaks claims 0.03 percent false positives, the lowest of any AI detector, and markets a second figure seven times higher. What both numbers leave out.

By the Undetected.ai team

August 2026 · 7 min read

The Humanizer

Try:
Tone
Strength:

2 free runs a day, up to 200 words each. We save your run so you can get back to it, and a delete button appears with the result. See privacy.

AI-pattern score

This is our own AI-pattern score, measured here on sentence rhythm, template phrases, vocabulary variety and passive voice. It is not a GPTZero, Turnitin, Originality.ai, Copyleaks or ZeroGPT result, and it does not predict one. Worth knowing: we also ask the rewrite to vary sentence length, drop template phrases and prefer the active voice, so some of the drop is built in. Read the two panels below, not just the number.

Before ·

After ·

·

Copyleaks publishes a false positive rate of 0.03 percent on English text, meaning it says it classifies 99.97 percent of genuinely human writing as human. That is the lowest published figure of any mainstream AI detector. It is also a vendor number, generated on data Copyleaks selected, and the company markets a second, different figure of 0.2 percent elsewhere. No independent test of Copyleaks false positives has been published.

That gap between what the number says and what it can support is the whole story here, so it is worth taking slowly. Below is where the figure comes from, what it does and does not cover, and how it compares against detectors that have been independently measured.

Where the 0.03 percent figure comes from

Copyleaks publishes a per-language accuracy table on its AI detector page. Read on August 4, 2026, the English row reads 99.97 percent for human text and 99.20 percent for AI text. Those two numbers measure different things and people mix them up constantly.

The 99.97 percent human figure is the one that matters for false positives. It says that out of every 10,000 human-written passages Copyleaks tested, three came back labelled as AI. Invert it and you get a 0.03 percent false positive rate. The 99.20 percent AI figure is the opposite direction: how much machine text it caught, where the miss becomes a false negative.

Copyleaks also states "over 99% accuracy" in its headline marketing and points to third-party studies for support. Separately, it has promoted a 0.2 percent false positive rate, roughly 1 in 500. Both figures are Copyleaks figures, and they differ from each other by almost seven times. That alone should tell you how much precision to read into either.

FigureWhat it claimsWhere it appears
0.03% false positives99.97% of human English text classed humanPer-language accuracy table on the AI detector page
0.2% false positivesRoughly 1 in 500 human documents flaggedCopyleaks marketing material
Over 99% accuracyCombined detection performanceHeadline claim on the product page
No independent figureNothing published by an outside teamNot available as of August 2026

How high is the Copyleaks false positive rate really?

Honestly, nobody outside Copyleaks knows. That is not a dodge, it is the actual state of the evidence, and it separates Copyleaks from several competitors.

GPTZero and Originality.ai have both been measured by an outside team. A 2025 University of Chicago Booth working paper built a corpus of 1,992 human texts written before 2020, paired them with 1,992 AI texts across several genres and lengths, and stress-tested detectors on short passages and on text run through humanizers. Both tools held at or below 1 percent false positives on medium and longer passages. Pangram came out at essentially zero on the same material. Copyleaks was not part of that comparison, so its claim sits unchecked.

The reasonable position is somewhere between trust and dismissal. Copyleaks is a serious product with a real research team, and the 0.03 percent figure is not invented. But every vendor number in this category is a best case: measured on text the vendor chose, at a confidence threshold the vendor chose. When outside teams have tested other vendors' best cases, the independent result has usually been somewhat worse. We keep the full comparison of AI detector false positive rates, vendor claims against independent measurement, on a separate page.

What the 0.03 percent figure does not cover

Three limits matter more than the decimal places.

Length. Detectors read statistical regularity, and a short passage does not contain much of it, so verdicts on short text swing hard in both directions. Copyleaks bills in credits where one credit covers 250 words, which nudges people toward scanning small chunks. A reading on a single paragraph is far less reliable than a reading on a full document, and no published accuracy figure from any vendor is measured on paragraph-length text.

Who wrote it. This is the largest effect in the entire field. Stanford researchers ran seven detectors over TOEFL essays by non-native English speakers and found 61.3 percent of that genuinely human writing classified as AI-generated, while native-speaker essays came back nearly clean. Copyleaks specifically markets low false positive rates on non-native English text and has published on the subject, which is a real point in its favor. It is still a claim the company makes about itself.

Which reading you get. Copyleaks returns a percentage of text it believes is AI-generated rather than a simple verdict, and people read that number as a confidence score, which it is not. We covered what the figure actually represents in our guide to the Copyleaks AI content score.

How Copyleaks compares to the detectors that have been measured

DetectorClaimed false positive rateIndependently measured
Copyleaks0.03% on English, also marketed as 0.2%None published
GPTZero0.9% at high confidenceAt or below 1% on medium and longer text
Originality.aiNot published as one figureAt or below 1% on medium and longer text
TurnitinUnder 1%, and 0.51% at document levelAbout 50% in a small 2023 press test
Pangram0.004% on academic essaysEssentially zero on medium and longer text

Turnitin is the cautionary row. Its own two published numbers already disagree, and the one independent test on record differs from both by a factor of roughly 100. That is not evidence that Copyleaks is wrong. It is evidence that the distance between a vendor's number and an outside measurement can be very large, and that Copyleaks has not yet been made to close that distance in public.

Is a low false positive rate the same as being safe?

No, and the arithmetic is the part people skip. A 0.03 percent rate applied to 10,000 submissions in a term still means three innocent people flagged. At the 1 percent that independent testing gives the better tools, the same 10,000 submissions produce 100. The rate feels small and the consequences do not scale down with it.

There is also a finding that cuts against the reflex to blame the software. Researchers publishing in Advances in Physiology Education in 2025 had 190 undergraduates write essays by hand, then ran them past four AI detectors and nine human graders. The detectors produced false positives on about 1.3 percent of essays. The human graders produced them on 5.0 percent. Detection tools were wrong, and the people checking by eye were wrong nearly four times as often. The study's practical conclusion was that agreement between several detectors, rather than any single reading, brings the false positive likelihood close to zero.

What to do if Copyleaks flags your writing

Lead with provenance rather than argument. A detector score is a statistical opinion about finished text, and you cannot rebut it with another statistical opinion. What does rebut it is a record of how the document was written: version history in Google Docs or Word, or a provenance tool such as Grammarly Authorship that replays composition from first paste to last keystroke. Those show process, which no detector reads.

Then ask the specific questions: which detector, on how many words, and at what threshold. A flag on 200 words is not a finding worth defending against. Our fuller walkthrough covers how to prove you did not use AI.

If the problem is recurring rather than a one-off, the cause is usually your writing style rather than any single scan. Plain, even, carefully edited prose scores as machine-like because that is literally the pattern being measured, which penalizes writers who were trained well. This bites hardest where volume is high and screening is automatic. Employers increasingly run applications through detectors, and anyone sending out a lot of applications, especially for the fully remote roles that draw hundreds of applicants each, is writing into exactly that filter. Varying sentence length and adding concrete specifics moves the measurement, because those are the properties being scored.

Does Copyleaks have a lower false positive rate than Turnitin?

On published claims, yes, by a wide margin: 0.03 percent against Turnitin's 0.51 percent at document level. In practice the comparison is close to meaningless, because the two tools are not used the same way. Turnitin cannot be run by the person being scanned, it refuses to score anything under 300 words, and it suppresses any reading between 1 and 19 percent, which conceals the band where its own false positives cluster. Copyleaks will scan anything you paste, at any length, and hand you a number.

That difference in accessibility matters more than the decimal point. A Copyleaks scan is something you can run yourself before submitting. A Turnitin score is something that happens to you. If Copyleaks is the checkpoint you are actually facing, our page on the best AI humanizer for Copyleaks covers what moves that score and what does not.

The short version

Copyleaks claims 0.03 percent false positives on English text, the lowest published figure in the category, and separately markets 0.2 percent. Neither has been checked by anyone outside the company. Both are best cases measured on long, clean text, and the real rate you experience depends far more on how much text you scan and who wrote it than on any number in a vendor table. Treat a single scan as one weak signal, not a verdict, and keep the record of how you wrote the thing.

Let Undetected.ai clear the flag for you

Paste your own text and watch our AI-pattern gauge sweep from the score on your draft to the score on the rewrite, meaning kept intact.

Make your next draft read like you wrote it

Paste your text and Undetected.ai rewrites the robotic patterns into natural prose, keeps your meaning, and scores the result on our own AI-pattern measure.

Meaning kept · Your own text rewritten · Saved to your history, delete any time

Humanize my text