Updated August 2026 · sourced from published research
Which AI is least detectable: Claude vs ChatGPT vs Gemini, and what actually moves an AI detector score
Search this question and you get confident percentages. Claude evades 24 percent of the time. ChatGPT is caught 92 percent of the time. Precise, quotable, and traceable to no published test.
Here is what the peer-reviewed work actually found, why the numbers everyone repeats are unreliable, and which variable is worth your attention instead.
2 free runs a day, up to 200 words each. We save your run so you can get back to it, and a delete button appears with the result. See privacy.
This is our own AI-pattern score, measured here on sentence rhythm, template phrases, vocabulary variety and passive voice. It is not a GPTZero, Turnitin, Originality.ai, Copyleaks or ZeroGPT result, and it does not predict one. Worth knowing: we also ask the rewrite to vary sentence length, drop template phrases and prefer the active voice, so some of the drop is built in. Read the two panels below, not just the number.
Before ·
After ·
Works on any model's output · Your own text rewritten · Saved to your history, delete any time
The short answer
No AI model is reliably undetectable, and the property that predicts detectability is model size rather than brand. The peer-reviewed AI Detectability Index examined 15 systems and found that larger models are consistently harder to detect than smaller ones, which puts Claude, GPT and Gemini in the same broad band and puts the small open models people run locally at the bottom, not the top. Claude is widely described as the least detectable of the three, and that belief is reasonable on the mechanics, but no independent test supports the specific percentages circulating. Those numbers come from companies selling humanizers. More usefully: switching between frontier models moves a detector score far less than changing how the text is written, because detectors score sentence-level predictability and rhythm, not authorship. The model you generate with is a starting position. The prose is the variable you control.
Last updated August 4, 2026. Sources on this page are peer-reviewed papers or vendor documentation read directly. Where a figure could not be verified we say so rather than repeating it.
The published evidence
What the peer-reviewed research actually found
Two academic benchmarks answer this question directly. Neither is quoted in the blog posts ranking for it, and both say something less convenient than "use Claude".
EMNLP 2023
The AI Detectability Index ranks models, and size wins
The Counter Turing Test paper set out to build exactly what people are searching for here: a quantifiable spectrum for ranking language models by how detectable they are. Its finding, in the authors' own words, is that across 15 contemporary models, "larger LLMs tend to have a higher ADI, indicating they are less detectable compared to smaller LLMs."
That is the closest thing to a real answer anyone has published. It also reframes the question. The axis is capability and scale, not which company's logo is on the chat window. A frontier model from any of the three major labs sits in the same region of that spectrum, and the small local model someone downloads specifically to stay private sits at the detectable end.
ACL 2024
RAID: the ranking is a moving target by design
RAID is the largest benchmark in this area: over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies, used to evaluate 8 open and 4 closed-source detectors. Its headline result is that current detectors are "easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models."
Read that last clause again, because it is the whole story. A detector is weak against models it has not been trained on. So any model is least detectable right up until it becomes popular enough that detector teams collect its output and retrain. "Which AI is least detectable" has an answer with a short shelf life, and choosing a tool on the strength of it means re-choosing every few months.
Model by model
Claude, ChatGPT, Gemini and the rest, on what can actually be checked
There is no honest "detection rate" column here, because no reproducible public test produces one. These are the properties that genuinely differ, and what each one implies.
| Model | Size class | What the ADI research predicts | Detector training coverage | What that means in practice |
|---|---|---|---|---|
| GPT-5 / GPT-4o (OpenAI) | Frontier, very large | Lower detectability than small models, by scale alone | Heaviest. Every detector on the market was built and tuned on GPT output first | The most written-about model, so it is also the most modelled by detectors |
| Claude (Anthropic) | Frontier, very large | Lower detectability than small models, by scale alone | Growing fast. Newer detector releases name Claude explicitly | Popularly believed to score lower. The belief is plausible on the research and unproven in public |
| Gemini (Google) | Frontier, very large | Lower detectability than small models, by scale alone | Substantial and rising | Sits in the same band as the other frontier models on everything anyone can actually verify |
| DeepSeek | Large, open weights | Similar to the frontier group at the top sizes | Lighter, but open weights let detector teams generate unlimited training data cheaply | Open weights cut both ways: less commercial attention, far easier for a detector to study |
| Llama (Meta) | Large to small, open weights | Small variants are the most detectable class in the research | Extensive. Llama output is a standard ingredient in detector training sets | The small quantized variants people run locally are the easiest text to flag, not the hardest |
| Mistral and other small open models | Small | Highest detectability of any group in the benchmark | Well covered in academic detection datasets | Running a smaller model to stay under the radar reverses the effect people expect |
The counterintuitive row is the last one. People assume a smaller, more obscure model is safer, and the research says the opposite.
Read this before you trust a percentage
Where the popular per-model numbers come from
Search "which AI is hardest to detect" and the first page is close to unanimous. Claude evades all five detectors 24 percent of the time. ChatGPT manages 4 percent. Detection rates of 87, 89 and 92 percent, one per model. The figures repeat across sites with slightly different wording, which reads like corroboration.
We went looking for the study. There isn't one. Every page carrying those numbers is published by a company that sells a humanizer or an AI detector, and none of them publishes a method, a sample size, a date, or the prompts used. Several cite each other. The precision is doing rhetorical work: 24 percent sounds measured in a way that "Claude seems a bit harder to detect" does not.
We are a humanizer vendor too, which is exactly why this matters. We would benefit from telling you that model choice is a solvable problem with a product-shaped answer. The accurate version is duller. Two academic groups have measured this properly, both found the effect tracks model scale, and neither produced a per-brand leaderboard you could act on.
If you want a number you can trust for your own writing, generate one. Paste the same 400-word passage into two or three free public detectors and record what they say, then do it again after a rewrite. That takes about ten minutes, costs nothing, and it is the only measurement on this page that describes your text rather than someone's marketing.
The other half of the question
Which models watermark their output, and why it does not change your answer
People choosing a model for detectability usually ask about detector scores and stop there. The second question worth asking is whether the model marks its own output, because a watermark is a definitive signal where a detector score is only a guess. The answer differs by vendor, and it is more interesting than the marketing on either side suggests.
OpenAI does not watermark ChatGPT text. It has built the capability: its provenance page states that its teams "have developed a text watermarking method that we continue to consider as we research alternatives", and that they "have prioritized launching audiovisual content provenance solutions" instead. The published reasons for holding it back are worth reading, because one of them is that the method "has the potential to disproportionately impact some groups. For example, it could stigmatize use of AI as a useful writing tool for non-native English speakers." That is the same group the Liang study found detectors misread 61.3 percent of the time.
Anthropic changed its answer in August 2026, and this is the newest fact in the whole subject. Having signed the EU AI Act's Article 50(2) transparency code, Anthropic now embeds an invisible watermark in text from Claude models "launched on or after August 2, 2026", worldwide rather than only in Europe, and across the API, Claude, Claude Code and the cloud platforms that resell it. In Anthropic's wording the model "weaves an imperceptible watermark directly into the text itself". Older Claude models are still being brought into scope. The catch is the same as Google's: Anthropic has not published the reader, so no third party can check a document for it today, and Anthropic says a detected mark "is not fully conclusive" about who wrote the text anyway. The full breakdown is on the Claude AI detector page.
Google is the exception. SynthID "embeds digital watermarks directly into AI-generated images, audio, text or video", and the text version runs in production on Gemini. DeepMind open sourced the method in October 2024, which is where most write-ups go wrong: open sourcing lets other developers watermark their own models with their own keys, and it does not let anyone verify Google's production Gemini output, because that needs Google's key.
So watermarking still should not move your model choice, though the reason has shifted. Two of the three major models now mark their text, and neither mark can be checked by your instructor, your client or any detector they are likely to run, because verification needs the vendor. ChatGPT leaves nothing to find at all. Every judgement anyone makes about your writing is still being made statistically, from the shape of the prose. We take the whole provenance stack apart in whether ChatGPT watermarks its text, including why the sites advertising a watermark detector cannot be reading one.
The variable that does move
Why changing models barely helps, and what does
Detectors do not look for a signature. They score two statistical properties of finished text. Perplexity measures how predictable each word is given the words before it. Burstiness measures how much that predictability varies across the document. Model output tends to be low on both: every word is a safe choice, and it stays a safe choice from the first line to the last. Human writing swings, because people change their minds mid-paragraph, over-explain one point and skip another, and drop in a short sentence for no structural reason.
Every frontier model produces text that is low-perplexity and low-burstiness by default, because that is what instruction tuning optimizes for. That is why the gap between them is narrow, and why switching from one to another moves a score by a little when it moves it at all. You are choosing between six shades of the same property.
Rewriting changes the property itself. A rewrite that genuinely varies sentence length, breaks the uniform paragraph shape, and replaces the model's default connective tissue moves the reading substantially, because it changes what the detector is measuring rather than which system produced the input. That holds whether the draft came from Claude, GPT, Gemini or DeepSeek, which is why our tool does not ask you which one you used.
The corollary is worth stating: prompting a model to "write like a human" is a weak version of the same idea and produces weak results. The model is still sampling from the same distribution, just with a stylistic instruction attached. Our breakdown of prompts that make AI writing sound human covers what those instructions can and cannot achieve.
Being straight with you
The thing no model choice and no rewrite touches
Everything above concerns detectors that guess from finished text. A second category of tool does not guess at all. Provenance features record how a document was written while you write it, and no choice of model and no rewrite has any effect on them.
Grammarly Authorship labels text as typed, pasted or AI-generated and lets a reader replay the writing process from the first paste to the last keystroke. GPTZero Writing Reports do something similar through Google Docs. Google Docs version history and Word revision tracking have quietly done a weaker version of it for years. These record rather than estimate, so there is nothing statistical to move.
The limits are real too. Authorship only records forward from the moment it is switched on, and only inside Google Docs or Microsoft Word, so it cannot be applied to a document retroactively. Turnitin now has a provenance feature of its own: the Clarity Writing Report records pasting activity, replays a timeline of the document being written, and surfaces AI chat interactions. It only applies inside Clarity assignments, so most Turnitin submissions are still scored the ordinary way, from the finished text alone. Where Clarity is switched on, no rewriting tool of any kind reaches it. We go through what Turnitin does and does not measure in our guide to how accurate the Turnitin AI detector is.
The practical version: if your work will be judged on how it was written rather than how it reads, model choice is not the lever and neither is a humanizer. Draft in the document, keep your notes and your version history, and let the record show the work. We go through what that evidence looks like in how to prove you did not use AI.
Worth holding in mind
The detectors doing the ranking are not that reliable either
Any per-model ranking inherits the error rate of the detector that produced it. On real-world text, independent testing puts the major detectors somewhere around 85 to 90 percent accuracy, with false positives in the 8 to 15 percent range. A study published in Patterns found that more than 60 percent of TOEFL essays written by non-native English speakers were misclassified as AI, against near-perfect accuracy on essays by US students.
Vendors are franker about this than their marketing suggests. Grammarly's own page states that "no AI detector is 100% accurate" and that "AI detection should never be used as a standalone verification method." Originality.ai tells customers to use detection as one signal rather than a final decision. GPTZero publishes that its accuracy is highest at document level and degrades at paragraph and then sentence level.
So a claim that model A is detected 89 percent of the time and model B 87 percent is quoting a two-point difference from instruments with error bars several times that wide. Even if the underlying test existed, the gap would not survive it. Our comparison of which AI detector is most accurate goes through each one's published claims and where they hold up.
Good questions
Questions people ask about AI models and detection
Which AI is least detectable?
No model is reliably undetectable, and the honest answer is that scale predicts detectability better than brand. The peer-reviewed AI Detectability Index found that larger models are consistently harder to detect than smaller ones across 15 systems. Among the frontier models, Claude, GPT and Gemini sit close together, and any gap between them shifts every time a detector retrains.
Is Claude less detectable than ChatGPT?
Possibly, but nobody has published a test you can check. The claim is repeated everywhere with confident percentages attached, and those percentages come from humanizer vendors marketing their own products rather than from independent testing. Claude does write with more varied sentence length on default settings, which is the property detectors read, so the belief is reasonable. It is still a belief.
Is Claude detectable by Turnitin?
Yes. Turnitin scores finished text and does not attempt to identify which model produced it, so Claude output is scored on the same statistical properties as anything else. Turnitin also refuses to score submissions under 300 words and suppresses any reading between 1 and 19 percent, which is why short Claude drafts often come back with an asterisk rather than a number.
Does switching AI models help avoid detection?
Barely, and it is the wrong variable to spend effort on. Switching from one frontier model to another moves a detector score far less than changing how the text is written, because detectors measure sentence-level predictability and rhythm rather than authorship fingerprints. Rewriting the prose changes what the detector actually measures. Changing the vendor mostly does not.
Why do AI detectors flag some models more than others?
Two reasons stack. Smaller models produce more predictable, lower-variance text, which is exactly the signal detectors score. And detectors are trained on the output they can collect most easily, so a widely used model is more thoroughly modelled. The RAID benchmark found detectors are "easily fooled by ... unseen generative models", meaning coverage matters as much as the model itself.
Can AI detectors tell which AI wrote something?
Mainstream detectors do not claim to. They return a probability that text is machine-generated, not an attribution to a specific model. Research into model attribution exists, but nothing in the commercial detectors students and editors actually face reports "this was Claude" or "this was GPT". Treat any tool claiming otherwise with suspicion.
Which AI writes the most human-sounding text?
Frontier models with higher default variation in sentence length and vocabulary read as more human, and Claude is the one most often named for this. That is a readability judgment rather than a detection result. A model that reads well to you can still score high, because the detector is measuring statistical regularity you cannot see by reading.
Does using a less detectable AI mean I do not need a humanizer?
No, and that is worth saying plainly even though we sell one. Picking a frontier model over a small one starts you from a better position, but every frontier model still gets flagged regularly on untouched output. The variable you control after generation is the prose itself, which is what a rewrite changes.
Go deeper
Related reading
Is Claude less detectable than ChatGPT?
The head-to-head, and why the confident percentages fall apart.
Can AI detectors detect Claude?
What happens when Claude output meets each of the five major detectors.
Which AI detector is most accurate?
Published claims against independent results, detector by detector.
Best AI humanizer overall
The full tool comparison, with prices we read ourselves.
Stop picking models, fix the prose
Whichever model wrote your draft, the rewrite is the part you control. Paste 400 words of the real thing and watch what happens to the reading.
Works on any model's output · Your own text rewritten · Saved to your history, delete any time