Independent · no humanizer or detector company pays us · prices re-checked daily
Methodology v1.0
Detectors / GPTZero accuracy

Detectors · primary sources

Is GPTZero accurate?

On plain AI text and on human writing, mostly yes. On humanized or lightly edited AI text, often no. GPTZero rarely flags human writing in independent tests, and catches most unedited ChatGPT-style text. But studies from 2025 and 2026 found it missed half or more of AI text after a humanizer rewrote it, and one peer-reviewed study found it missed every fully AI paper at its strict threshold.

Accuracy for an AI detector is two numbers, not one: how often it wrongly flags human writing (false positives) and how often it misses AI text (false negatives). A detector can look excellent on one and poor on the other. Below are GPTZero's own claims, then what independent researchers measured, each with its date, because detectors change model versions often.

What GPTZero claims

  • Its homepage headline is 99% accuracy. Its FAQ cites the RAID benchmark: 95.7% of AI texts caught while wrongly flagging 1% of human texts.
  • On non-native English writers, GPTZero says it cut its false-positive rate on TOEFL essays to 1.1%, and in October 2026 published its own test where its latest model labelled all 914 English-learner essays in a school dataset as human.
  • It markets "paraphraser" and "AI bypasser" detection. But its own FAQ page, live on the same day we read the homepage, says the classifier is not trained to identify AI text after it has been heavily modified. Both statements can't be the whole story.
  • GPTZero also says results shouldn't be used to punish anyone or as the final verdict.

These are vendor claims. Here's what others measured.

What independent tests found

Study (date) What was tested GPTZero result
RAID benchmark paper, ACL 2024 (GPTZero queried Jan 2024) 11 AI models, many domains, set to wrongly flag 5% of human text 66.5% of AI text caught overall: very high on chat models (ChatGPT 99.4%), low on older base models. Most robust of the commercial detectors to tricks like look-alike characters.
RAID live leaderboard (read Oct 8, 2026; entry undated) Same benchmark, later GPTZero submission 95.7% caught at a 1% false-positive rate, falling to 92.9% when adversarial attacks are included. 18th-19th of about 45 entries.
Liang et al., Stanford, Patterns, Jul 2023 (detectors from Mar 2023) 91 TOEFL essays by non-native writers, 7 detectors including GPTZero Across the seven detectors, 61.3% of the human TOEFL essays were wrongly flagged on average. An old version of GPTZero; its current claims are above.
Jabarian & Imas, NBER working paper, Sep 2025 Several genres and AI models, including text run through the StealthGPT humanizer A "secondary tier" behind Pangram. After StealthGPT, GPTZero missed around half or more of the AI text in most settings.
Van Vlasselaer et al., Int. J. for Educational Integrity, Jun 2026 (peer-reviewed; text from May 2025) 160 English papers: human, fully AI, hybrid, and AI rewritten by a GPT-4o prompt No human paper wrongly flagged, but no fully AI paper caught at the strict level either; 0% on hybrid papers and 2.5% on the rewritten ones.
Karr et al., Notre Dame, arXiv preprint, Aug 2026 Published research abstracts, before and after rewriting with Undetectable AI 0% of pre-ChatGPT abstracts flagged. After Undetectable AI rewriting, GPTZero missed 96.1% of the AI-labelled texts.
DAMAGE paper, Jan 2025 (written by Pangram Labs, a competing detector) 19 humanizer tools on academic text, 5% false-positive setting 99.7% of plain AI text caught, but 60.0% after humanizing.

A pattern runs through all of them: GPTZero is careful about accusing humans, which is the right way round for a tool used on students. The cost is that rewritten or edited AI text often gets through.

So should you trust a GPTZero score?

  • If it says "human": that's weak evidence. Humanized and lightly edited AI text regularly scores human.
  • If it says "AI" on unedited text: stronger, especially on long English prose, which is where GPTZero says it works best.
  • If it flags your own writing: it happens less often than with older detectors, but it does happen. GPTZero itself says its result shouldn't decide a case. Keep your drafts and version history; they're better evidence than any score.
  • Short texts (a paragraph or two) are less reliable for every detector.

What GPTZero costs

GPTZero's pricing page showed paid plans only when we read it; the homepage lets you scan up to 10,000 characters without an account. Current plans, yearly vs monthly prices and the word allowance are on our GPTZero pricing page.

Our own free-detector test starts in November 2026: the same human-written and AI-written texts through GPTZero and the other free detectors, with every text and score published. Until then, this page reports other people's measurements. How we test.

Sources (all read Oct 8, 2026)

  1. GPTZero homepage, FAQ, /faq and /technology pages (undated): gptzero.me
  2. GPTZero blog, "GPTZero has no bias against ESL writers", Oct 5, 2026: gptzero.me/news
  3. Dugan et al., "RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors", ACL 2024, arXiv 2405.07940 (v2 Jun 2024)
  4. RAID leaderboard: raid-bench.xyz/leaderboard
  5. Liang, Yuksekgonul, Mao, Wu & Zou, "GPT detectors are biased against non-native English writers", Patterns 4(7), Jul 2023
  6. Jabarian & Imas, "Artificial Writing and Automated Detection", NBER Working Paper 34223, Sep 2025
  7. Van Vlasselaer, Van Droogenbroeck & Spruyt, International Journal for Educational Integrity 22:16, Jun 29, 2026
  8. Karr, Khvatskii, Hua & Chawla, "Why AI Detection Fails for Academic Integrity", arXiv 2608.11256, Aug 2026
  9. Masrour, Emi & Spero (Pangram Labs), "DAMAGE: Detecting Adversarially Modified AI Generated Text", arXiv 2501.03437, Jan 2025