Undetected.ai
All posts
Comparisons

Is Claude Less Detectable Than ChatGPT? What the Evidence Shows

Probably a little, and nobody has published a test you can check. What the peer-reviewed benchmarks found, where the confident per-model percentages actually come from, and why the model you draft in matters less than the rewrite.

By the Undetected.ai team

August 2026 · 8 min read

The Humanizer

Try:
Tone
Strength:

2 free runs a day, up to 200 words each. We save your run so you can get back to it, and a delete button appears with the result. See privacy.

AI-pattern score

This is our own AI-pattern score, measured here on sentence rhythm, template phrases, vocabulary variety and passive voice. It is not a GPTZero, Turnitin, Originality.ai, Copyleaks or ZeroGPT result, and it does not predict one. Worth knowing: we also ask the rewrite to vary sentence length, drop template phrases and prefer the active voice, so some of the drop is built in. Read the two panels below, not just the number.

Before ·

After ·

·

Probably a little, and nobody has published a test you can check. Claude's default output varies sentence length and vocabulary more than GPT's, which is exactly the property AI detectors score, so the belief is mechanically reasonable. The confident percentages attached to it are not: every one traces back to a company selling a humanizer or a detector. The peer-reviewed work that has looked at this found the effect tracks model size rather than brand, which puts Claude and GPT in the same band and makes the question less useful than it looks.

This comes up constantly, and usually from someone deciding which chat window to draft in. It is a reasonable instinct and a poor investment of effort. Below is what is actually known, what is merely repeated, and where the real difference sits.

Is Claude less detectable than ChatGPT?

On the mechanics, there is a plausible case. Detectors score two things: perplexity, meaning how predictable each word is given what came before, and burstiness, meaning how much that predictability swings across a document. Machine text is low on both. Claude's default register runs longer and more varied at the sentence level than GPT-4o's, which tends toward uniform paragraph blocks with a recognizable connective rhythm. More variation reads as more human to a classifier measuring variation.

On the evidence, there is almost nothing. No university group, no detector vendor and no independent lab has published a reproducible head-to-head with a stated method, sample size and date. What exists instead is a cluster of blog posts quoting figures like "Claude evaded all five detectors 24 percent of the time" against "ChatGPT managed 4 percent." Follow any of them to the source and you land on another humanizer's marketing page, or on a sibling post at the same company.

We sell a humanizer, so we would gain from telling you this is settled and simple. It isn't. The defensible statement is that Claude may start you from a marginally better position, and that no one can tell you by how much.

What the peer-reviewed research actually found

Two academic benchmarks address this directly, and neither produces the brand leaderboard people are looking for.

The Counter Turing Test paper, presented at EMNLP 2023, built the AI Detectability Index specifically to rank language models by detectability. Examining 15 contemporary models, the authors concluded that "larger LLMs tend to have a higher ADI, indicating they are less detectable compared to smaller LLMs." Scale is the axis, not the logo. Claude, GPT and Gemini are all frontier-scale systems and sit together in that region, while the small open models people run locally to stay private are the most detectable class in the study.

The RAID benchmark, presented at ACL 2024, is the largest dataset in the field: over 6 million generations across 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies, used to test 12 detectors. Its finding is that detectors are "easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models."

That last phrase does the damage. A detector performs badly against models it was not trained on, which means any model's advantage is temporary by construction. Detectors were built on GPT output first because GPT was what there was. As Claude's share has grown, detector teams have collected Claude output and retrained on it. Whatever gap existed in 2024 is smaller now, and the direction of travel is fixed.

Claim you will seeWhere it comes fromWhat can actually be said
Claude evades detection 24% of the time, ChatGPT 4%Humanizer vendor blogs, no method publishedUnverifiable. Do not plan around it
Detection rates of 87%, 89%, 92% by modelSame cluster of sites, quoting each otherTwo-point gaps from instruments with much wider error bars
Larger models are harder to detect than smaller onesEMNLP 2023, 15 models testedSupported by peer review, and the useful finding
Detectors fail on models they have not seenACL 2024, 6M+ generationsSupported, and it makes any ranking temporary
Claude writes with more sentence variation by defaultObservable in the outputFair as a description, not as a detection rate

Why does Claude set off fewer AI detectors than ChatGPT?

Where a difference does show up, two things explain it. The first is stylistic: Claude's instruction tuning produces prose with more variation in sentence length and a less formulaic paragraph shape, and variation is the burstiness signal in the detector's favor. The second is exposure: GPT output has been the training substrate for AI detection since the field started, so detectors model it in more detail than anything else.

Neither is a property of Claude being fundamentally different in kind. Both are contingent, and the second is actively closing. It is the difference between a model being genuinely harder to identify and a detector simply having had less practice.

Is Claude detectable by Turnitin?

Yes. Turnitin scores finished text and makes no attempt to identify which model produced it, so a Claude draft is judged on the same statistical properties as anything else that arrives in the submission portal. Turnitin also refuses to score submissions under 300 words and suppresses any reading between 1 and 19 percent, showing an asterisk instead, because it treats readings in that band as too unreliable to report.

That last detail matters more than model choice for most students. It means a short Claude draft frequently comes back with no number at all, which people read as clearance when it is really an abstention. We go through the mechanics in how Turnitin detects AI.

Does switching AI models help avoid detection?

Barely, and it is the wrong variable. Moving between frontier models changes a detector reading by a little, because all of them produce low-perplexity, evenly paced prose by default. You are choosing between shades of the same property.

Rewriting changes the property. Genuinely varying sentence length, breaking the uniform paragraph shape and replacing the model's default connectives moves a reading substantially, because it alters what the detector measures rather than which system fed it. That holds regardless of origin, which is the argument for spending your effort after generation rather than before it. The full comparison across models lives on our page on which AI is least detectable.

Prompting a model to "write like a human" is a diluted version of the same idea. The model still samples from the same distribution with a style note attached, and detector readings move accordingly: a bit, inconsistently.

Can AI detectors tell which AI wrote something?

Mainstream detectors do not claim to. They return a probability that a passage is machine-generated, not an attribution to a named model. Model attribution is an active research area, but nothing in the commercial tools students and editors actually encounter reports "this was Claude." If a product tells you it can, treat that as a marketing claim until it shows you the method.

This is worth knowing because it undercuts the premise of the question. If detectors cannot tell models apart, then "which model is least detectable" is really asking which model happens to sit furthest from the classifier's decision boundary on average. That is a statistical accident of training data, not a durable property you can buy into.

How reliable are the detectors doing the ranking?

Any per-model comparison inherits the error rate of the detector producing it. On real-world text, independent testing puts the major detectors around 85 to 90 percent accuracy with false positives in the 8 to 15 percent range. A study published in Patterns found more than 60 percent of TOEFL essays written by non-native English speakers were misclassified as AI, against near-perfect accuracy on essays by US students.

The vendors say so themselves when you read past the homepage. Grammarly's own page states that "no AI detector is 100% accurate" and that "AI detection should never be used as a standalone verification method." Originality.ai tells customers to treat detection as one signal rather than a final decision. GPTZero publishes that accuracy is highest at document level and falls off at paragraph and sentence level.

So a claimed two-point gap between models is noise from instruments with error bars several times wider. Our run-through of which AI detector is most accurate covers each one's published claims against what holds up.

What no model choice fixes

Everything above concerns tools that guess from finished text. A separate category does not guess. Grammarly Authorship labels text as typed, pasted or AI-generated and lets a reader replay the writing process from the first paste to the last keystroke. GPTZero Writing Reports do something comparable through Google Docs. Version history in Google Docs and revision tracking in Word have done a cruder version for years.

These record rather than estimate, so there is no statistical property for a model swap or a rewrite to move. They have real limits: Authorship only records forward from activation, only inside Google Docs or Word, and cannot be applied to a finished document after the fact. But within those limits, no choice of model matters at all. If your work will be judged on how it was written, draft in the document and keep the trail. We cover what that evidence looks like in how to prove you did not use AI.

Who is actually asking this, and what they should ask instead

Two groups land on this question and they need different answers. Students are usually one submission away from a consequence, and for them the model is close to irrelevant next to the length threshold, the suppression band and their institution's policy. Marketers and content teams are asking something else entirely, because Google does not run an AI detector on your draft and has said repeatedly that it rewards helpful content regardless of how it was produced.

For the second group the detection question is often a proxy for a harder one. If you are publishing at volume, the constraint is rarely whether a classifier flags the prose; it is whether the piece was worth writing, which is a keyword research problem rather than a detection one. Our take on the actual ranking risk is in does Google penalize AI content.

The practical answer

Draft in whichever model writes best for your purpose. If that is Claude, the small detectability edge is a bonus rather than a reason, and it is shrinking. Then treat the draft as a draft: read it aloud, break the even paragraphs, cut the connectives that give it away, and check it against two or three public detectors before it matters. Ten minutes of that beats any amount of model shopping, and it produces a number about your text rather than someone's marketing.

One late development cuts against Claude for the first time. In August 2026 Anthropic started embedding an invisible watermark in text from Claude models launched on or after August 2, worldwide, while OpenAI still ships no text watermark at all. Neither fact is checkable by anyone outside the vendor today, because Anthropic has not released the reader, so it does not change a detector score. It does mean the model with the marginally better statistical profile is now also the one carrying a provenance signal, which is a genuinely new consideration if you are choosing on this basis. The detail is on the Claude AI detector page.

If you want the rewrite done properly rather than by hand, that is what our AI humanizer is for, and it does not ask which model wrote the input because it does not need to.

Let Undetected.ai clear the flag for you

Paste your own text and watch our AI-pattern gauge sweep from the score on your draft to the score on the rewrite, meaning kept intact.

Make your next draft read like you wrote it

Paste your text and Undetected.ai rewrites the robotic patterns into natural prose, keeps your meaning, and scores the result on our own AI-pattern measure.

Meaning kept · Your own text rewritten · Saved to your history, delete any time

Humanize my text