How Accurate Are AI Detectors? Understanding AI Detector Reliability


Reading time
3 mins
 How Accurate Are AI Detectors? Understanding AI Detector Reliability

Bottom line: AI detectors often flag human-written academic text as AI-generated because they assess how predictable and formulaic the text is.

What is an AI detector?

AI detectors are tools that claim to detect whether a piece of text was written by humans or by AI. They do this by inspecting the text for certain statistical qualities, measuring how “predictable” the text is. AI-generated text is almost always uniform in tone and style, so AI detectors look for this.

How does an AI detector work

AI detectors measure 2 key properties in any text:

  1. Perplexity: This is how predictable or “smooth” your word choice is. Human writing has higher perplexity than AI writing.
  2. Burstiness: This is how uniform your sentences are in length and structure. Humans usually write a mix of short and long sentences while AI produces sentences with near-uniform length and complexity.

Are AI detectors accurate? AI detector reliability

From the time they emerged in 2022, the accuracy of AI detectors has been constantly challenged. The creators of ChatGPT, OpenAI, rolled back their own AI detector because of accuracy issues. MIT Sloan Teaching and Learning Technologies says flat out that “AI detection software is far from foolproof—in fact, it has high error rates”.

Controversies around Pangram

In May 2026, Pangram, whose AI detector has rapidly gained popularity, published a case study titled “AI is Writing Prize-Winning Fiction”, claiming that one of the stories winning the 2026 Commonwealth Short Story Prize, The Serpent in the Grove by Jamir Nazir, was 100% AI-generated. In June 2026, Razmi Farook, Director-General of the Commonwealth Foundation, issued a statement that they have thoroughly investigated the allegation, and examined working drafts, time-stamped documents, and notes provided by all their writers. The Commonwealth Foundation concluded that the story was genuinely human-written.

Pangram also made headlines when its CEO, Max Spero, accused on social media 3 Wall Street Journal columnists of producing AI-generated articles. WSJ’s editorial page editor James Taranto strongly hit back in a column titled “The ‘AI Detector’ as Defamation Machine,” pointing out that Pangram itself moderates its claims about its tool but its CEO goes on AI witch hunts on social media.   

AI detector accuracy for academic writing

AI detectors so far are more likely to throw up false-positives for academic writing. As mentioned earlier, academic writing is low-perplexity; research papers are written using precise terminology and controlled vocabulary. A patient “presents with diarrhea” and doesn’t “show up with the runs”. Burstiness is comparatively lower in research writing too. Exclamations and asides just don’t belong in research papers. Have you ever read: “A total of 76 patients were lost to follow up. Sigh. Anyway, we’re pretty stoked that 512 patients completed treatment and we’ve got their data to show you. Hang in there!”

Student writing is arguably a lot less predictable and not so well structured than a dissertation or journal article. Yet, many universities have explicitly discontinued AI checkers or advised faculty against using such tools. For example, Boston University advises professors: “Be very cautious with accusations of GenAI misuse—all detection tools are highly fallible, both with respect to false positives and false negatives, despite the marketing claims of companies that sell these products.”

Similarly, Johns Hopkins University warns instructors about relying on AI detection scores, “If misconduct is suspected, it is highly recommended that instructors speak to students before making any accusations.”

Most major academic journals and publishers permit the use of generative AI, with proper disclosure. AI for language editing, translation, and even drafting text is widely acceptable as long as the final paper is free from hallucinations like fake citations, and flat, meaningless writing. AI use in research images and figures is still controversial, and of course, journals expect a proper AI disclosure statement that outlines which model you used, version, and which tasks AI performed.

AI detector evaluation: How to choose an accurate AI detector

  • Choose a tool trained on an academic corpus. General-purpose tools are not suitable for academic writing.
  • Choose a tool that gives you a sentence-wise breakdown; a single overall score tells you less.
  • Check the false-positive rate (i.e., how often the tool flags a human-written paper as AI).
  • Remember that many AI detectors are trained on older models of ChatGPT, Claude, etc. In one study, AI detection was substantially more accurate for ChatGPT 3.5 than for GPT-4. The vendor should disclose which LLMs and models their tool has been trained to detect.
  • Remember that free online third-party AI detectors may retain your paper as training data. Opt for a tool that your institution holds a license for. Else, read the tools data retention and data security documentation before you input anything.
  • Do not purchase or use an AI humanizer. These often distort your meaning and alter valid technical terms.

What you should look for in an AI detector review: Choosing a reliable AI detector

Reviews of AI detectors are often geared towards generating more sales for that tool, rather than solving researchers’ pain points. Always consider the following before acting on any review:

  • Where it is published: be more suspicious of reviews posted on the vendor’s own website or social media pages
  • Who authored it: an employee of the vendor company, a member of the public, or an academic?
  • Whether it is sponsored or paid for: reputable publications like blogs and magazines indicate whether something is a sponsored post or advertorial.
  • What was actually tested: an AI detector will not perform the same on a student essay and on a systematic review.

Practical tips for using AI detection tools

Here’s how to use an AI detector in the best and wisest way:

  • Remember that the Methods section is likely to have higher AI scores than other sections because sampling, recruitment, data collection, and analysis all need to be described using conventional phrasing and controlled vocabulary. Perplexity and burstiness are both lowest in this section.
  • If your paper has to follow rigorous guidelines like PRISMA and CONSORT, your score will be higher. These guidelines require precise, specific wording (again, low perplexity). For example, “risk of bias” in a systematic review cannot become “chance of prejudice”.
  • Do not run the detector on your acknowledgments, funding statement, or reference list.
  • Always inform your co-authors if you’re uploading any parts of the paper into an AI detector, regardless of whether it’s licensed or free.
  • Treat an AI score as an indicator only. Remember that different tools will produce different scores (I’ve even got different scores from the same tool on different days).
  • Never introduce grammatical errors, typos, or unconventional wording to lower your AI score. You’re just creating another set of problems for yourself by doing so.
  • AI detectors do not check hallucinations or fabricated citations. You need a separate workflow to verify these.

Author

Marisha Fonseca

An editor at heart and perfectionist by disposition, providing solutions for journals, publishers, and universities in areas like alt-text writing and publication consultancy.

See more from Marisha Fonseca

Found this useful?

If so, share it with your fellow researchers


Related post

Related Reading