How Do AI Detectors Work? What They Measure
AI detectors measure how predictable and how even your writing is. What perplexity and burstiness really are, and why a 90 percent score is misread.
By the Undetected.ai team
August 2026 · 8 min read
2 free runs a day, up to 200 words each. We save your run so you can get back to it, and a delete button appears with the result. See privacy.
This is our own AI-pattern score, measured here on sentence rhythm, template phrases, vocabulary variety and passive voice. It is not a GPTZero, Turnitin, Originality.ai, Copyleaks or ZeroGPT result, and it does not predict one. Worth knowing: we also ask the rewrite to vary sentence length, drop template phrases and prefer the active voice, so some of the drop is built in. Read the two panels below, not just the number.
Before ·
After ·
AI detectors work by measuring how predictable your writing is. A language model reads the text and calculates, word by word, how likely each choice was given everything before it. Machine-written prose is unusually predictable and unusually even in rhythm, so a classifier trained on millions of examples of both kinds of writing returns a probability that the passage looks machine written. It is a statistical guess about style. No detector reads a hidden mark, and none of them can tell you which model produced the text.
That last sentence is where most of the confusion in this subject starts, so it is worth being precise about what is actually happening inside these tools before looking at what their numbers mean.
What perplexity and burstiness actually measure
Two properties do most of the work in every detector on the market.
Perplexity is a measure of surprise. Feed a sentence to a language model and it will assign a probability to each next word. If the words that actually appear are the ones the model would have picked, perplexity is low. Human writing tends to have higher perplexity because people make odd choices: a word slightly off register, a piece of jargon from their own field, a construction that a model averaging over its training data would rank lower. A generated draft is, almost by definition, made of the likely choices.
Burstiness is variance in that predictability across a document, and in sentence structure alongside it. Human writing is bursty. A twelve word sentence lands next to a forty word one. A dense technical paragraph is followed by a throwaway aside. Generated prose is flatter: sentence after sentence in the same band, each built to a similar shape. That evenness is the single most reliable tell, and it is the property that survives most rewriting attempts.
Modern commercial detectors do not compute these two numbers and stop. They feed features like these into a supervised classifier trained on large paired corpora of human and machine text, and the classifier learns whatever separates the two in that data. The vocabulary of perplexity and burstiness survives because it describes what the model is picking up on, not because there is a formula anyone is reading off.
What a 90 percent AI score really means
Almost everyone reads this number wrong, including people making decisions with it.
GPTZero documents the answer in its own API reference. Its classifier returns a document_classification field with three possible values, HUMAN_ONLY, MIXED and AI_ONLY, alongside a probability for each. On what that probability means, GPTZero writes: "The class probability corresponding to the predicted class can be interpreted as the chance that the detector is correct in its classification. I.e. 90% means that 90% of the time on similar documents our detector is correct in the prediction it makes."
So a 90 percent score is not a claim that 90 percent of your document was written by AI. It is a statement about how often the detector is right when it feels this confident. Those are completely different claims, and the second one is much weaker than the first. A student told that a detector returned 90 percent typically hears an accusation about nine tenths of their essay. What the tool said was closer to "when I am this sure, I turn out to be right about nine times in ten", which also means it is wrong roughly one time in ten.
| What you see | What it literally means | What people assume it means |
|---|---|---|
| "90% AI" | How often this detector is correct on documents it scores this way | 90% of the words were AI generated |
| Sentence highlighting | Where in the document the classifier's confidence concentrated | These exact sentences were copied from a chatbot |
| Confidence category | A tuned band. GPTZero says that at high confidence, 99.1% of human articles are classified human and 98.4% of AI articles as AI | A guarantee attached to this particular document |
| A named model in the marketing | That model's output was in scope during training | The tool identified which model wrote the text |
The fourth row is the one that matters most for anyone comparing tools. GPTZero states that it "works robustly across a range of AI language models, including but not limited to ChatGPT, GPT-5, GPT-4, GPT-3, Gemini, Claude, and AI services based on those models". That is a coverage claim. It is not model attribution, and no commercial classifier does model attribution, which is why every product branded as a Claude AI detector or a GPT detector is running the same general classifier behind a different landing page.
Why detectors flag human writing
If the thing being measured is predictability and evenness, then any human who writes predictably and evenly will score badly. That is not a bug being fixed in the next release. It follows directly from the method.
The people it lands on are consistent. Non-native English speakers write with a smaller working vocabulary and more standard constructions, which is exactly the profile of low perplexity. Liang and colleagues, publishing in Patterns in 2023, ran seven detectors across TOEFL essays written by human students and found an average misclassification rate of 61.3 percent. Technical and procedural writing has the same problem for a different reason: structured, repetitive prose scores as machine-like because it is machine-like in form. GPTZero notes that its classifier "can sometimes flag other machine-generated or highly procedural text as AI-generated".
Length matters more than most people realize. Every detector is less reliable on short passages, because there is less variance to measure. GPTZero says accuracy at document level is greater than at paragraph level, which is greater than at sentence level. Turnitin refuses to score anything under 300 words at all. If someone is judging a paragraph, the tool they used was operating in its weakest mode. We collected the measured error rates, and traced the widely repeated ones that lead nowhere, on the AI detector false positive rate page.
Why editing changes the score
Because the classifier was never trained on edited text. GPTZero puts this plainly in its own FAQ: "Our classifier is not trained to identify AI-generated text after it has been heavily modified after generation."
That single sentence is the honest mechanism behind this entire product category, ours included. Rewriting does not defeat a detector through a trick. It moves the text out of the distribution the classifier learned to recognize, by restoring the variance and the odd choices that generated prose lacks. It is also why rewriting a sentence at a time rarely helps: swapping synonyms leaves the rhythm and the paragraph architecture intact, and those are what is being measured.
It is worth saying what this does not buy you. A rewrite is a language operation. It does not check facts, so an invented citation survives it in more convincing prose. It does not add the specific detail that makes writing worth reading, and it does not settle anything about provenance, which is a separate question with a separate kind of evidence.
What AI detectors cannot do
Three limits are structural rather than temporary.
They cannot identify a model, for the reason above. They cannot prove authorship, because they read finished text and infer backwards from style, which is why version history and dated drafts carry more weight than any score in an actual dispute. And they cannot be audited by the person being scored, which is the sharpest issue in academic use: Turnitin's AI indicator is visible to instructors and administrators, not to the student whose work is being judged.
Detector vendors are often more careful about this than the people quoting them. GPTZero recommends its results "should not be used to punish students" and describes the classifier as a way "to flag situations in which a conversation can be started". Turnitin says its score "does not make a determination of misconduct". The caution gets stripped out somewhere between the documentation and the meeting.
The same gap is opening up in hiring, where scanning cover letters through a detector has become common enough to generate its own search traffic, despite there being no published evidence on false positive rates for that kind of short, formulaic document. Since many employers now put candidates through automated first-round screening interviews regardless, the detector pass on a cover letter is doing very little work for the amount of unfairness it can cause.
Do AI detectors read a hidden watermark?
Not currently, with one change worth tracking. Detectors are statistical classifiers; a watermark would be a definitive signal, and a definitive signal returns a yes or a no rather than a percentage. If you are being shown a percentage, no watermark is being read.
The change is that Anthropic began embedding an invisible watermark in Claude text in August 2026 under the EU AI Act transparency code, and Google already watermarks Gemini output with SynthID. Neither is readable by a third-party detector: verification needs the vendor's key, and Anthropic has said the technical documentation for detecting its marks is still to come. OpenAI, meanwhile, built a text watermarking method and has not shipped it, which we went through in whether ChatGPT watermarks its text.
The short version
AI detectors calculate how predictable and how even your writing is, then compare that profile against what they learned from millions of human and machine documents. The output is a probability about style, and the percentage refers to how often the detector is right rather than how much of your document was generated. They flag people who write plainly, they get less reliable the shorter the text gets, they cannot name a model, and their own vendors say the results should not be used as proof. Knowing which question a number is answering is most of what you need to argue with one.
Let Undetected.ai clear the flag for you
Paste your own text and watch our AI-pattern gauge sweep from the score on your draft to the score on the rewrite, meaning kept intact.