All articles

AI content detectors in 2026: what actually works, and what does not

A short field report on what AI content detection can and cannot do in 2026, written by a team that sells it.

LT
Lacewing Technologies
·

Every few months someone forwards us the same screenshot: a paragraph they wrote themselves, flagged as 98% AI. Then, in the same week, someone else sends a piece written entirely by a model that sailed through the same checker untouched. Both things are true at once, and understanding why is the whole story of AI content detection in 2026.

We build detection tools. TextSight scores written text and Truthring examines recorded speech, so we have a commercial interest in people believing detectors work. We would rather be straight with you about what they do and do not do, because overselling this is how the whole category loses trust.

Where accuracy actually sits

The honest headline number in 2026 is this: no general-purpose detector is reliably above about 85% accuracy across every model it might encounter. The strongest tools land near 82%, and detectors miss somewhere between 15% and 30% of machine-written text depending on which model produced it.

That is not a scandal. A screening tool that catches four out of five is genuinely useful. It becomes a problem the moment someone treats the output as a verdict rather than a signal.

The false positive problem is not evenly distributed

Across tested tools, false positive rates run between roughly 3% and 12%. But the average hides who actually pays for it.

Non-native English writers get flagged at rates between 5% and 19%, against 1% to 6% for native speakers. The reason is mechanical rather than malicious: detectors largely measure how predictable text is. Someone writing carefully in a second language tends to use plainer vocabulary, tighter sentence structures and fewer idioms, which is exactly what low-perplexity, machine-like text looks like to a statistical model. Technical and academic writing gets caught the same way, for the same reason.

If you run a university, a hiring process or a marketplace, this is the number that should shape your policy. A 10% false positive rate applied to thousands of submissions is a lot of wrongly accused people, and they will not be randomly distributed.

Light editing beats detection

Detection accuracy drops 20% to 30% when someone restructures sentences and swaps terminology. Add genuine original insight and rewrite properly, and accuracy falls below 50%, the point at which you are guessing.

This has an uncomfortable implication for anyone hoping detection will settle the question of authorship: the people most motivated to evade detection are the easiest to miss, and the people writing honestly in a plain style are the most likely to be flagged. The tool is weakest exactly where you most want it to be strong.

What the platforms actually do

It is worth separating two questions people constantly merge: can this be detected, and does anyone penalise it.

Google’s position has been consistent. It evaluates content quality, not whether a machine helped write it. Its ranking systems look at expertise, experience, authority and trust signals, and at whether the page is accurate and useful. There is no ranking penalty for AI assistance as such; there is a penalty for thin, unhelpful, mass-produced pages, which is a different thing and always has been.

Academic institutions have largely landed in the same place: detection as a screening step that triggers a human review, never as evidence on its own. Legal and compliance teams treat scores as probabilistic signals, because that is all they are.

So what is detection good for?

Used properly, it is a triage tool. It tells you where to look, not what to conclude. That makes it genuinely valuable in a few situations:

  • Volume screening. When 4,000 submissions arrive and you can only read 200 closely, a score tells you which 200.
  • Consistency checks. A writer whose style suddenly changes across a body of work is worth a conversation, whatever the tool says.
  • Vendor and marketplace quality. If you are paying for original writing and receiving generated filler, detection plus a human read catches it quickly.
  • Voice and audio. Synthetic speech is a different problem from synthetic text, and the stakes (fraud, impersonation, fabricated evidence) are higher. This is where the category is heading.

The rule we work to: a detection score should never be the last step in a decision that affects someone. It should be the first step in a process that ends with a person.

How to buy a detector without getting burned

  1. Ask for the false positive rate, not the accuracy rate. Accuracy is easy to inflate by tuning aggressively. Any vendor who cannot tell you their false positive rate on human-written text has not measured it.
  2. Ask what it was tested against. A detector tuned on last year’s models degrades quietly as new ones ship. Ask when it was last re-measured and on what.
  3. Check how it handles non-native writing. If they have not tested this, assume the equity problem above applies to you.
  4. Insist on a three-outcome result. “Likely synthetic”, “likely human” and “unclear” is honest. A single percentage with no uncertainty band is marketing.
  5. Look for something a third party can verify. A reference code or reproducible report matters the moment a result is disputed.

Where we stand

We publish false positive rates alongside detection rates, we return an “unclear” verdict when the evidence is thin, and for Truthring we are not quoting accuracy figures at all until they are measured and published, which is why it is still marked pre-launch rather than shipped with numbers we made up.

Detection in 2026 is a useful instrument with real limits. Treat it that way and it will serve you well. Treat it as proof and it will eventually cost someone their job, their grade or their reputation, and they will not have deserved it.

Lacewing Technologies builds AI detection and content systems, including TextSight.ai for written text and Truthring for recorded speech. If you need detection built into your own product or process, get in touch.

Share LinkedIn X

Need this built into your product, not bought off a shelf?

We build detection, classification and content-quality systems for teams who need them wired into an existing workflow. Tell us what you are trying to catch and we will tell you whether detection is even the right instrument.