There are now more AI detectors than there are reasons to use one. Most of the “best detector” lists you will find are affiliate pages that rank whoever pays the most, quote a 99% accuracy figure straight off a vendor’s homepage, and never mention what happens when the tool is wrong.
This is our attempt at a more useful version. We have ordered these by what they are actually good at, said plainly where each one falls down, and included the numbers that matter rather than the ones that market well.
Disclosure, up front: two of the ten tools on this list are ours. TextSight.ai and Truthring are built by Lacewing Technologies. We have marked them clearly, listed their weaknesses alongside everyone else’s, and you should discount our opinion of them accordingly. The other eight we have no stake in.
Before the list: the three numbers that decide everything
Almost every buying mistake in this category comes from looking at the wrong metric. Three things matter, in this order.
1. False positive rate, not accuracy
Accuracy is trivially easy to inflate: tune the model to call almost everything AI and your detection rate goes through the roof. The number that costs you is the false positive rate: how often it flags genuine human writing as machine-made.
Across independently tested tools, false positives run somewhere between 3% and 12%. The best of them are now under 1%. The worst are far higher than their marketing suggests. If a vendor cannot tell you their false positive rate, they have not measured it, and you are the experiment.
2. How it treats non-native English writers
This is the part the industry keeps quiet about. Non-native English writers get falsely flagged at rates between 5% and 19%, against 1% to 6% for native speakers. The mechanism is not prejudice, it is arithmetic: most detectors score how statistically predictable text is, and someone writing carefully in a second language uses plainer vocabulary and tighter structures, which reads to the model exactly like machine output. Technical and scientific writing trips the same wire.
If you are screening students, job applicants or freelancers at any scale, this is the single most important number in this article.
3. Whether it survives light editing
Detection accuracy falls 20% to 30% when someone restructures sentences and swaps a few terms. Rewrite properly and add real insight, and it drops below 50%, which is coin-flip territory. Any tool claiming it catches everything is either not testing against edited text or not telling you the truth.
Keep those three in mind and the list below sorts itself out.
1. Pangram Labs: the most accurate general detector
Pangram has quietly become the tool to beat on raw accuracy. In a 30-tool comparison run under identical conditions, using roughly 1,000 words each from GPT-4o, Gemini 2.0 and Claude 3.7 Sonnet plus three human-written control texts, Pangram scored 9/9 on AI text and 3/3 on human text. Only one other tool matched it.
Its stated false positive rate is under 1%, which is the lowest credible figure in the category, and it explains its verdicts at phrase level rather than handing back a single percentage. For teachers, it groups submissions by class, which is a small thing that saves a lot of clicking.
Where it falls down: it needs at least 50 words, so it is no use on short answers or social posts. The free tier is four credits a day. And the comparison above was published by Pangram itself, which is worth remembering, because vendor-run benchmarks tend to be kind to the vendor.
Price: free tier with daily credits, around $20/month for premium.
Use it if: accuracy on unedited text matters more than anything else.
2. TextSight.ai: detection and rewriting in one place (ours)
Every other tool on this list tells you there is a problem and stops. TextSight is built around the assumption that the next thing you want to do is fix it.
It scores a piece of writing, highlights the specific sentences driving the score, and then rewrites those sentences in place with three intensity settings: light, balanced, or maximum. For anyone editing content at volume, that removes a step: you are not copying flagged text into a second tool and pasting it back.
Two deliberate design choices are worth explaining, because they are unusual.
The first is that it returns a 0–100 authenticity score rather than a binary verdict. A binary answer implies a confidence that the underlying science does not support. A score tells you where you sit and how far you have to move, which is a more honest description of what the model actually knows.
The second is that we do not publish a single headline accuracy number. Competitors advertise 99.41% and 99.98%. Those figures come from internal tests on curated datasets, and they collapse the moment you hand the tool edited text or writing from a model it has not seen. We publish the methodology instead. It is a worse marketing decision and a more defensible one.
Where it falls down: we are not the most accurate detector on this list. Pangram and Copyleaks both beat us on unedited machine text, and we would rather say so than pretend otherwise. If detection is all you need and you never rewrite anything, one of them is a better fit. TextSight makes most sense when detection and editing are the same job.
Price: free tier, then $9.99/month for Starter, $19.99 for Pro with unlimited scans and API access, $39.99 for Business with five seats. Annual billing takes 25% off.
Use it if: you are producing or editing content and need to act on the result, not just measure it.
3. Truthring: for voice, not text (ours, pre-launch)
Everything else on this list reads writing. Truthring listens to audio, and it is here because the fraud has moved.
Synthetic voice is now the higher-stakes problem: a cloned voice on a phone call authorising a transfer, a fabricated voice note used as evidence, a family member who is not actually the person calling. Text detection is largely an academic-integrity and content-quality issue. Voice detection is a fraud issue, and the money involved is not comparable.
Truthring examines a recording across three things. The recording chain, meaning the microphone and room signature a real capture leaves behind and a rendered file does not; the generator signature, which can point at the family of system that produced it; and prosody under stress, which is how a voice behaves at the awkward edges of natural conversation, where synthesis is still weakest. It returns one of three verdicts (likely synthetic, likely human, or unclear) plus a reference code someone else can check.
Where it falls down: it is not shipped yet. The analysis pipeline is documented in full and the product is in pre-launch, but we have not published measured accuracy figures, and we are not going to quote any until we have. That is the honest state of it, and if you need voice detection in production today you will have to look elsewhere for now.
Price: not yet published.
Use it if: your risk is impersonation and fraud over audio rather than authorship over text, and you can wait for the measured numbers.
4. Copyleaks: the enterprise choice
Copyleaks matched Pangram’s perfect score in that same 30-tool test: 9/9 on AI text, 3/3 on human. What separates it is everything wrapped around the detection: customisable detection profiles, proper document handling, plagiarism checking in the same product, and the compliance paperwork large organisations need before they can buy anything.
It handles long documents better than most, which matters if you are checking dissertations or full manuscripts rather than blog posts.
Where it falls down: it is slow. Roughly a minute to process a 200-word test in one comparison, which is fine for a batch job overnight and irritating if you are checking things one at a time. The free allowance is five scans.
Price: from around $16.99/month for a credit bundle.
Use it if: you are buying for an organisation and need audit trails, profiles and plagiarism in one contract.
5. Originality.ai: built for publishers and agencies
Originality was designed for a specific buyer: the person paying freelancers for original writing who needs to know what they are getting. It scans in bulk, keeps a history against team members, and combines AI detection with plagiarism and readability in one report.
In independent testing it caught 7 of 9 AI samples and correctly cleared all three human texts. Respectable, though short of the top two.
Where it falls down: it advertises 99.41% accuracy, which is a laboratory figure, not a field one. It is also credit-based rather than unlimited, so heavy scanning gets expensive faster than the headline price suggests.
Price: around $14.95/month base, credit top-ups on top.
Use it if: you commission writing at volume and need a paper trail per contributor.
6. GPTZero: the one built for teachers
GPTZero came out of the education world and still fits it best. It runs a layered detection model, shows sentence-level highlighting, gives writing feedback alongside the score, and has an API and a Chrome extension. In testing it caught 7 of 9 AI samples with no false positives on the human control texts.
Its most useful feature is also the one people complain about: it hedges. Rather than declaring a document AI or human, it will often return “mixed” and show you which passages drove that. That is scientifically the correct behaviour and pedagogically the right one too, because it forces a conversation instead of an accusation.
Where it falls down: if you want a clean yes or no to paste into a disciplinary form, this is not your tool. It is deliberately uncomfortable to weaponise.
Price: free tier around 10,000 words a month, premium from about $15/month.
Use it if: you teach, and you want evidence for a conversation rather than a verdict.
7. Winston AI: the best-connected one
Winston’s advantage is that it goes where your work already is. It integrates with Zapier and Google Classroom, has browser extensions, and includes OCR, which means it will read a photographed or scanned page, not just pasted text. For anyone dealing with handwritten or printed submissions, that is a genuine differentiator rather than a feature-list entry.
It also bundles plagiarism checking, so for many teams it replaces two subscriptions.
Where it falls down: it advertises 99.98% accuracy, which should be read as a marketing number rather than a measurement. It needs an account before you can test it and a 500-character minimum, so quick spot checks are awkward.
Price: 14-day trial, then from around $12/month on annual billing.
Use it if: you need detection inside an existing workflow, or you deal with scanned and photographed documents.
8. Turnitin: the institutional default
Turnitin is on this list because of where it sits rather than how it performs. It is embedded in the learning management systems universities already run, which means for a great many institutions it is not a purchasing decision at all. It is simply what appears when a student submits work.
That reach is its strength and its risk. A tool used on millions of submissions makes even a low false positive rate into a large absolute number of wrongly flagged students, disproportionately those writing in a second language. To its credit, Turnitin’s own guidance is explicit that its score is an indicator requiring human review, not proof of misconduct, guidance that is regularly ignored by the people using it.
Where it falls down: individuals cannot buy it, you cannot easily audit it, and its ubiquity means institutional habits form around it faster than the evidence justifies.
Price: institutional licensing only.
Use it if: you are a university and it is already in your stack, and pair it with a written policy that forbids acting on the score alone.
9. Sapling: the one to build on
Sapling is aimed at developers and support teams rather than editors. It offers sentence-level scoring, fast processing, an API designed to be embedded in other products, and it updates promptly when new models appear, including the open-weight ones many detectors ignore until much later.
In testing it caught 6 of 9 AI samples and correctly cleared all three human texts, so it is more conservative than the leaders: it misses more, but it does not accuse innocent writing.
Where it falls down: the free tier is tiny at 2,000 characters, and the product is not really designed for someone who just wants to check one essay.
Price: free tier, around $25/month for Pro.
Use it if: you are putting detection inside your own product or support pipeline.
10. ZeroGPT: the free one, with a warning
ZeroGPT is the tool most people meet first, because it requires no signup and costs nothing. It supports multiple languages and even runs through messaging bots.
Its results, though, are the least reliable on this list. In independent testing it caught 6 of 9 AI samples but cleared only 1 of 3 human texts, and in one test it flagged Claude-generated content with a 58% false positive reading. As a rough first look it is fine. As the basis for any decision affecting a person, it is not.
Where it falls down: false positives, heavy advertising on the free tier, and no meaningful methodology published.
Price: free up to 15,000 characters, around $10/month for premium.
Use it if: you want a zero-commitment sanity check and nothing rests on the answer.
Why detectors disagree so violently with each other
Run the same paragraph through four tools and you will often get four different answers. That is not a bug in any one of them; it is a consequence of what they are all measuring.
Nearly every text detector is scoring some version of predictability. Language models pick likely next words, so machine text tends to sit in a narrow band of statistical expectation, low perplexity in the jargon, with unusually even sentence-to-sentence variation. Human writing wanders. It has odd word choices, sentences that run long and then stop short, unnecessary asides.
The problem is that plenty of human writing does not wander. Legal drafting does not wander. A lab report does not wander. A careful writer working in a second language does not wander, because wandering is what you do when you are relaxed in a language. All of that reads as machine-like to a model that only knows how to measure surprise.
Different vendors weight these signals differently, train on different corpora, and update at different speeds. So disagreement between them is not a sign that one is broken. It is a sign that you are asking a question none of them can answer with certainty, which is exactly why the ones that return “unclear” are being more honest than the ones that never do.
Run your own test in twenty minutes
Do not take anyone’s benchmark, including this article’s. Detection performance depends heavily on the kind of text you actually handle, and a tool that wins on essays can lose badly on product descriptions. Testing it yourself is quick.
- Collect ten pieces of writing you are certain a person wrote. Ideally your own team’s work, from before generative tools were in common use, and including at least two from non-native English speakers if that reflects your users.
- Generate ten pieces on the same topics using two or three different models, not one. Detectors perform very differently depending on which system produced the text.
- Edit five of the generated ten the way a real person would: restructure some sentences, swap terminology, add a genuine opinion. This is the batch that separates marketing from reality.
- Run all twenty through each shortlisted tool and record the raw score, not your impression of it.
- Count two things only: how many human pieces were wrongly flagged, and how many edited AI pieces slipped through. Ignore the unedited AI results. Every tool on this list handles those reasonably well, so they tell you nothing useful.
Twenty samples is not statistically rigorous, but it will separate the tools that work on your material from the ones that do not, and it takes an afternoon rather than a procurement cycle.
Four mistakes that cost people real money
Treating a score as evidence. The single most expensive error. Universities have reversed decisions and employers have settled complaints over accusations built on a percentage. Every vendor’s own documentation says the output is an indicator. Put that in writing in your policy before you need it.
Buying on the headline accuracy figure. 99.41% and 99.98% are internal numbers from curated tests. They are not predictions about your content. The gap between advertised and observed performance is widest exactly where the stakes are highest.
Standardising on one tool forever. Detection quality drifts as models change, and it drifts silently. There is no alert when your detector stops recognising the newest system. Re-test every six months with the method above. It is cheaper than finding out during a dispute.
Screening people without telling them. Beyond the fairness question, undisclosed screening tends to fail the moment it is challenged. If you check submissions, say so, say what happens when something is flagged, and say who reviews it. The tools work better when people know the rules.
How to choose between them in about two minutes
Strip away the marketing and there are really only five buying situations.
- You need the most accurate read on unedited text. Pangram or Copyleaks. Nothing else is close on independent testing.
- You need to fix what you find. TextSight, because detection and rewriting sit in the same screen.
- You are screening people, not pages. GPTZero, because it hedges honestly and produces evidence for a conversation rather than a verdict.
- You are buying for an organisation. Copyleaks or Turnitin, depending on whether you need a contract or already have one.
- You are building detection into your own product. Sapling for text. Truthring for voice, once it ships.
Five questions to ask any vendor before you pay
- What is your false positive rate, and on what corpus? If the answer is a single accuracy percentage, they are answering a different question.
- When did you last re-measure, and against which models? A detector tuned on last year’s models decays quietly as new ones ship. There is no warning light.
- How does it perform on non-native English writing? If they have not tested this, assume the 5–19% figure applies to your users.
- Does it return an “unclear” verdict? A tool that is never uncertain is not being honest about what it knows.
- Can a third party verify a result? The moment a score is disputed, and it will be, you need something reproducible.
The rule we build to: a detection score should never be the last step in a decision that affects someone. It should be the first step in a process that ends with a person reading the work.
What this list will look like in a year
Three things are already visible.
Text detection is getting harder, not easier. Each generation of model writes with more variance, which is precisely the signal detectors rely on. The tools at the top of this list will stay near the top, but the gap between “caught” and “missed” will keep widening for anyone who edits their output even lightly.
Voice and video are where the urgency has moved. A wrongly flagged essay is unfair. A cloned voice authorising a payment is theft. Budgets follow harm, and the harm is now loudest in audio.
Provenance will eventually beat detection. The durable answer is not guessing after the fact but signing at the point of capture: cryptographic provenance attached when a camera, microphone or editor creates a file. That infrastructure is being built now. Detection is the bridge until it arrives, and bridges are temporary by design.
Until then, use these tools for what they are: instruments that narrow the search, not machines that deliver verdicts. Every one of them, ours included, is wrong often enough that a human has to stay in the loop.
Lacewing Technologies builds AI detection and content systems. TextSight.ai for written text, Truthring for recorded speech — and builds detection into other companies’ products. If that is what you need, tell us what you are working on.
Sources and caveats: comparative pass rates come from a 30-tool test published by Pangram Labs, run on roughly 1,000-word samples from GPT-4o, Gemini 2.0 and Claude 3.7 Sonnet against three human control texts — a vendor-run benchmark, and read as such. False positive ranges and the non-native English figures come from independent 2026 analyses of detector reliability. Prices are as advertised in September 2026 and change often; check the vendor before buying.
