All articles
,

The false positive problem: who AI detectors wrongly accuse, and why

Detector errors are not random. They concentrate on second-language writers, neurodivergent writers and anyone writing in a constrained register. The evidence, the arithmetic, the court record, and a procedure that holds up.

LT
Lacewing Technologies
·

The uncomfortable fact about AI detection is that its errors are not distributed randomly. They land, over and over, on the same people: writers whose second language is English, writers who are neurodivergent, and writers whose prose is simply plain. Understanding why is the difference between a workable policy and a machine for producing injustices.

We sell a detector. We would rather sell it to people who understand this, because the alternative is selling it to people who will eventually use it to ruin someone’s year.

The study that changed the conversation

In July 2023, Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou at Stanford published GPT detectors are biased against non-native English writers in Patterns. They ran seven detectors (Originality.AI, Sapling, Crossplag, GPTZero and ZeroGPT among them) over two sets of unambiguously human essays: 91 TOEFL essays written by Chinese speakers, and 88 essays written by American eighth-graders.

61.22%average misclassification rate on human-written TOEFL essays
19.78%of those essays flagged as AI by all seven detectors at once
97.80%flagged by at least one detector

Performance on the American eighth-grade essays was near perfect. Same tools, same day, same task, and a catastrophic error rate on one population against almost none on the other.

Then they did the thing that turned a correlation into a mechanism. They asked ChatGPT to improve the word choice in the TOEFL essays, leaving the ideas and structure alone, and re-ran the detectors. Misclassification fell from 61.22% to 11.77%. Unanimous flagging fell from eighteen essays to one.

Then they ran it backwards: they simplified the native-speaker eighth-grade essays to resemble non-native phrasing. Misclassification rose from 5.19% to 56.65%.

The detectors were not measuring authorship. They were measuring lexical richness, and lexical richness is a proxy for fluency, not for honesty.

That is the whole finding, and it explains every downstream failure. Generation-one detectors keyed on perplexity. Plain, high-frequency vocabulary produces low perplexity. So does machine output. A second-language writer using a smaller working vocabulary lands in the same statistical region as GPT-3.5, and the tool cannot tell the difference because it was never measuring the difference.

Everyone this catches

Non-native English writers are the documented case, but the mechanism is general. It catches anyone whose writing is more predictable than the median.

  • Autistic and other neurodivergent writers. Consistent structure, literal phrasing and low stylistic variation are exactly the features that read as low burstiness. Moira Olmsted, an autistic student at Central Methodist University, was given a zero on a reading summary; the accusation was withdrawn when she disputed it, but she was told a future flag would be treated as intentional. The AI Incident Database catalogues detection tools misidentifying neurodivergent and second-language students’ work as a recognised pattern.
  • Anyone who used a writing assistant. Marley Stevens at the University of North Georgia was placed on academic probation for a year after Turnitin flagged an essay she had written with Grammarly. Grammarly’s whole function is to reduce the statistical irregularity of prose. It is a false-positive generator by construction.
  • Anyone who wrote in another language first. Weber-Wulff and colleagues found that running genuinely human text through Google Translate or DeepL cost roughly twenty percentage points of detector accuracy. Machine translation flattens exactly the idiosyncrasies detectors rely on.
  • Anyone writing in a constrained register. Technical documentation, legal drafting, lab reports, regulatory filings and clinical notes are all genres where deviation is a defect. They are also genres detectors flag heavily, because the conventions that make the writing good are the conventions that make it predictable.

The arithmetic institutions ignore

Here is the calculation that made Vanderbilt University disable Turnitin’s AI detector on 16 August 2023, less than four months after it launched.

Vanderbilt submitted roughly 75,000 papers to Turnitin in 2022. Turnitin’s own claimed false positive rate at the time was 1%. That is around 750 papers per year falsely flagged at a single institution, and Vanderbilt would have no way to know which 750.

They also noted, pointedly, that Turnitin would not explain what “patterns common in AI writing” meant, which made it impossible for a student to mount a defence or for a faculty member to evaluate the evidence.

The base-rate problem is the one nearly every institutional deployment gets wrong. A 1% false positive rate sounds like a rounding error until you multiply it by volume. Turnitin reported reviewing more than 200 million papers in the first year of its AI detector. One percent of 200 million is two million.

What the vendors actually publish

Turnitin discloses more than most competitors, and the disclosures are more revealing than the headline claim.

DisclosureWhat it actually means
Document-level false positive rate below 1%Only for documents already scoring 20% AI or above. The rate outside that band is not published.
Sentence-level false positive rate around 4%Roughly one in twenty-five highlighted sentences is human-written. Sentence highlights are the thing instructors actually read.
About 54% of false-positive sentences sit next to genuine AI textThe highlights are least reliable at exactly the boundaries a reviewer scrutinises hardest.
Scores below 20% hidden since July 2024The vendor concluded low-range scores were unreliable enough that showing them was a liability.
Minimum 300 words of qualifying proseShort submissions cannot be scored at all. Any tool that scores a paragraph is guessing.

That fourth row is the most honest thing any detection vendor has done. Deciding not to display a number is an admission that the number was not good enough to act on, and it was the number institutions were acting on most often.

The independent measurements

Bloomberg Businessweek ran the cleanest test we know of, published 18 October 2024. Jackie Davalos and Leon Yin obtained college application essays via public records request and ran them through GPTZero and Copyleaks.

On 500 essays submitted to Texas A and M in summer 2022, written before ChatGPT existed and therefore human by definition, 1 to 2% were falsely flagged as AI, in some cases with near-total stated confidence. On 305 essays from summer 2023, around 9% were flagged, though those cannot be assumed human.

One to two percent is a dramatic improvement over the 61% Liang measured on TOEFL essays a year earlier. It is still, at national admissions volumes, tens of thousands of false accusations.

What changed by 2026

We should be as fair to detection as we are critical of it, because the 2023 numbers are now widely quoted as if they were current, and they are not.

A 2026 study in the International Journal for Educational Integrity by Van Vlasselaer, Van Droogenbroeck and Spruyt at Vrije Universiteit Brussel tested four current tools (Pangram, Turnitin, GPTZero and Copyleaks) against long-form documents of four thousand words or more, plus a real-world corpus of over eleven hundred master’s theses. On the fully human-written arm, all four tools produced zero document-level false positives.

The failure mode has inverted. In the same study, several tools missed most of the machine-written papers entirely, and performance on humanised text was worse still. The dominant 2026 error is the false negative, not the false accusation.

Two caveats keep this from being a clean acquittal. That study used long documents. Four thousand words gives a detector an enormous amount of signal, and the false-positive problem was always concentrated in short text. And it did not test multilingual or second-language cohorts, which is precisely the population where the original bias was found. Until someone repeats Liang’s TOEFL experiment against 2026 detectors and publishes it, the honest position is that document-level false positives on long English text appear to be largely solved, and that nothing has been demonstrated about short text or non-native writers.

What the courts have said

The leading case is Haishan Yang v. Neprash, decided in the US District Court for the District of Minnesota on 31 October 2025.

Yang, an international PhD student, was expelled from the University of Minnesota over alleged AI use on a preliminary exam. The evidence included faculty judgement that the answers did not match his voice, side-by-side comparisons with ChatGPT output the professors generated themselves, and a GPTZero score of 89% on one answer and 19% on another.

The court upheld the university. The reasoning is what matters for anyone writing policy: substantive due process requires showing the institution acted in bad faith or departed substantially from accepted academic norms, and courts defer to academic bodies where a decision is reasoned, evidence-based and within professional judgement, explicitly including novel areas like AI detection.

Two lessons. First, the detector score was not what carried the case; the corroborating comparison was. Second, litigation is not the safeguard. Institutional procedure is. If your process is sound, courts will not second-guess it. If your process is a detector score and a form letter, courts still may not second-guess it, which should worry you more, not less.

The governance gap

The Center for Democracy and Technology surveyed US middle and high school teachers across the 2023 to 2024 school year. Around 68% regularly use AI detection tools. Around 64% report that students have faced negative consequences over AI use or accusations, up sixteen percentage points year on year.

Only 28% had received any guidance on how to respond to a suspected case.

That is the actual problem. It is not that detectors are used. It is that they are used by people who have been given a number and no procedure, at institutions that never wrote one. And the survey found licensed special-education teachers using detection tools at a notably higher rate than their colleagues, against a student population that uses generative AI more than its peers, a compounding risk for the group least able to absorb it.

A procedure that holds up

If you are responsible for one of these decisions, this is the minimum defensible shape.

  • The score opens an inquiry. It never closes one. No sanction should be issuable on detector output alone, and your written policy should say so explicitly, because that sentence is what protects your staff as much as your students.
  • Corroborate with process evidence. Version history, drafts, search history in the document, an oral conversation about the content. A student who wrote the work can discuss it. A student who did not, generally cannot, and that conversation is better evidence than any classifier.
  • Check the factual layer. Fabricated citations and confidently wrong specifics remain the most reliable signal available, and they require no tooling.
  • Disclose the score and its error rate to the accused. If you will not tell someone what the number was and how often it is wrong, you do not have a process, you have a verdict.
  • Set a length floor. Below 300 words, do not score at all.
  • Track your own outcomes. How many flags did you raise, how many survived investigation? If you are not measuring your own false positive rate, you are relying on the vendor’s — measured on their data, not yours.

Disclosure. Lacewing Technologies builds TextSight.ai, an AI text detector. Everything above applies to our product as much as anyone else’s. We would rather you deploy detection with a procedure around it than deploy it confidently and get someone expelled on a number.

The shape of the answer

Detection is not going to become a machine that reliably identifies cheating, because that is not what it measures and it never was. What it measures is statistical resemblance to machine output, and there will always be humans whose writing legitimately resembles machine output.

The workable version treats the score as a smoke alarm: a cheap, imperfect signal that something is worth looking at, attached to a process where a human does the looking. Alarms that go off when you make toast are annoying. Alarms wired directly to the fire brigade are a problem. The engineering question was never how to eliminate false alarms; it was what happens after one.

If you want the mechanics behind the scores, we covered how detectors actually work. If you are choosing a tool, we ranked the major detectors including our own, with weaknesses stated.

Sources and further reading

  • Liang, Yuksekgonul, Mao, Wu and Zou, GPT detectors are biased against non-native English writers, Patterns, July 2023
  • Weber-Wulff et al., Testing of detection tools for AI-generated text, International Journal for Educational Integrity, 2023
  • Van Vlasselaer, Van Droogenbroeck and Spruyt, International Journal for Educational Integrity, 2026
  • Davalos and Yin, Do AI Detectors Work?, Bloomberg Businessweek, 18 October 2024
  • Vanderbilt University, Guidance on AI detection and why we’re disabling Turnitin’s AI detector, 16 August 2023
  • Turnitin, published false positive rate guidance and AI writing detection model documentation
  • Center for Democracy and Technology, Up in the Air, teacher survey
  • Haishan Yang v. Neprash, US District Court, District of Minnesota, 31 October 2025
Share LinkedIn X

Need this built into your product, not bought off a shelf?

We build detection, classification and content-quality systems for teams who need them wired into an existing workflow. Tell us what you are trying to catch and we will tell you whether detection is even the right instrument.