AI detectors do not work, and using them is a decision with consequences
We ran human and machine text through five detectors. False positives were common enough to make any policy built on them unfair.

Part of AI writing in real work, not in demos
A detector accusing a student of using AI on an essay that student demonstrably wrote is not a rare failure reported by a handful of unlucky people. It is the predictable output of a category of software asked to do something the underlying method cannot support, deployed anyway because a score is administratively convenient and an actual investigation is not. We built a corpus and ran it through five detectors to find out how bad the problem actually is, rather than take either side's word for it.
The corpus, and why it had to include text older than the models
We assembled three kinds of writing. Clearly human text, some of it published years before GPT-3 existed so contamination was impossible — student essays from a decade ago, blog posts from the early 2010s, internal memos with dated file properties. Clearly generated text, produced with a plain prompt and no editing pass. And hybrid text, built by taking a generated draft and rewriting it the way an actual working person would: cutting a third of it, replacing the flattest sentences, reordering two paragraphs, correcting a claim that turned out to be wrong.
That third category matters most and gets studied least. Almost nobody writes with a model today by asking for a finished document and shipping it unread; the writers we tracked in our notes from testing the writing tools on live documents mostly edit rather than ship raw output. The real population of AI-touched text is the hybrid, and if a detector's accuracy claims come from a corpus of pure generated versus pure human, the number does not describe the writing anyone is actually trying to evaluate a policy against. We ran all three categories through five widely used detectors, scored each piece as flagged or not against the threshold each tool sets by default, and recorded the result without adjusting for what we expected to see.
False positives, and who they land on
Across the pre-2020 human writing — text that cannot possibly be AI-generated because it predates the models — every detector produced false positives. Not on the same pieces, and not at identical rates, but consistently enough that no detector cleared human writing at a rate we would call reliable. The pattern held worst on plain rather than stylistically distinctive writing: short declarative sentences, low use of idiom, correct but unremarkable grammar.
That pattern has a name in the research literature, and it is the single most damaging fact about how these tools get deployed: false positive rates rise for non-native English writers. The reason is not mysterious once you see the mechanism. Detectors score how predictable a sequence of tokens is — how close each word choice sits to the most statistically likely next word given what came before. Fluent native writing tends to be varied in its predictability, full of idiom and register shifts and the occasional odd turn of phrase a person reaches for without thinking. Writing produced by someone working carefully in a second language is often more uniformly correct and plain, because the writer reaches for the construction they know is safe rather than the one that sounds most like them — precisely the statistical signature the tools read as machine-generated.
Put a detector in front of an admissions office, a hiring pipeline or an academic integrity process, and the population most likely to be wrongly flagged is already facing the steepest disadvantage in that process. That is not a hypothetical edge case. It is the mechanism working exactly as designed, on exactly the writing it is worst at reading.
The hybrid case breaks the tools completely
The generated-and-then-edited text is where every detector we tested performed worst, and the failure is structural rather than a bug that improves with the next model update.
A detector measures statistical uniformity — the signature of text produced by the same process under the same conditions throughout. A person who takes a generated draft and genuinely edits it — cutting, restructuring, replacing weak sentences, correcting claims — breaks that uniformity by construction. The result is neither purely one thing nor purely the other in the way the detector's training data assumes text must be. Some detectors respond by flagging it as human because the smoothness is gone; others flag it as AI because enough of the original tokens survived the edit. Neither answer describes what actually happened, which is that a person wrote a document using a tool, the way a person writes a document using a spell-checker, and the question "did AI write this" does not have a binary answer to give.
This matters because the hybrid case is becoming the median case, not a corner case being stress-tested. Most people who use these tools use them the way we describe in our piece on where AI writing actually earns its place and the case for editing rather than drafting — as an editing pass on a person's own draft, not a replacement for having one. A detection regime built to sort "human" from "machine" as a clean binary measures a category that describes less and less of the writing it is asked to judge.
Why the confidence number on the screen is not a probability
What we found genuinely indefensible is not that detectors make mistakes — every classifier does — but how the mistake is presented. A detector reports a confidence percentage, that number looks like a probability, and it gets treated like one in the meeting where a decision gets made.
It is not a probability that a person or a model wrote the text. It is a measure of how well the text's statistical properties match a training distribution the vendor built from a limited and not fully disclosed sample. A high score means "this text's predictability profile resembles our generated-text examples more than our human-text examples" — a much weaker claim than "there is an X percent chance a machine wrote this," and the two are not close enough to round together. A weather forecast is validated against thousands of prior days with known outcomes and recalibrated continuously; a detector's score is validated against whatever corpus the vendor happened to build and recalibrated only when the vendor chooses to retrain. Nobody making a hiring or grading decision sees that context. They see a percentage and read it as evidence.
What the vendors themselves say, which contradicts how the tools are used
Read the documentation and terms of use for any major detector and a pattern repeats: the vendors hedge in language the institutions deploying their product routinely ignore. The formulations are close to identical across the category — scores should be used as one signal among several, the tool should not be the sole basis for an accusation or an academic integrity finding, human judgment should make the final call. That is not a minor footer disclaimer. It is a direct statement, from the people with the most complete knowledge of how their own tool fails, that it should not be used the way it is overwhelmingly being used — as a single number that ends a conversation instead of opening one.
A school that suspends a student on the strength of a detector score, or an editor who rejects a freelancer's pitch because a checker flagged it, is doing exactly what the vendor's documentation says not to do. The tool is not misused by accident. A single number is administratively convenient and a real investigation is not, and the hedge is easy to skip past in a terms page nobody reads before the product goes into a policy.
What to do instead if the question actually matters
If you genuinely need to know how a piece of writing was produced, not as a box to tick but because something real depends on the answer, the honest options do not run through a detector at all.
Ask for process, not a verdict on the finished text. Version history, whether from a document editor's revision log or simply drafts saved at different times, shows a trajectory a single generated-and-pasted document does not have. It is not proof by itself — a person can save fake intermediate drafts — but it is a different kind of evidence than a statistical guess about token predictability, and harder to fake convincingly than it is to generate text.
Ask the person to explain and extend what they wrote, on the spot, without the document in front of them. Someone who wrote it themselves can usually walk you through why a paragraph is ordered the way it is and defend a claim under a follow-up question. Someone who pasted in a generated draft they never fully absorbed often cannot, and that gap shows up in five minutes far more reliably than in a percentage score.
Where the stakes genuinely require certainty, design the assessment so the method of production cannot answer the question either way — an in-person component, a source the writer has to demonstrate they understood, a conversation about the work itself. That costs more than uploading a document to a checker, and it produces an answer that means something.
The honest position is not that detectors are slightly unreliable and should be weighted accordingly. It is that the number they produce cannot bear the weight it is being asked to carry in the places it is now installed, and an institution that has wired a detector score into a consequential decision should be able to explain, to the person on the wrong end of a false positive, exactly what evidence beyond that score justified the outcome. In most of the policies we looked at, there isn't one.
Questions people ask
- Can AI detectors be trusted to prove a student or employee used AI?
- No. Every detector produces false positives on genuinely human writing, and the rate rises sharply for non-native English speakers whose sentence patterns happen to resemble the statistically smooth output detectors are tuned to flag. A single score is not evidence.
- Why do AI detectors flag human writing as machine-generated?
- Detectors measure how predictable a sequence of words is, not who typed it. Writing that is plain, correct and low in surprising word choices scores as "predictable" whether a person wrote it carefully or a model produced it automatically, so competent, unadorned prose is exactly the kind of human writing most likely to be misread.
- Do detectors work on text that has been edited after being generated?
- Badly. A generated draft that a person then rewrites, trims and restructures is now the most common kind of text in real workplaces, and it is the case detectors handle worst, because editing breaks the statistical uniformity the tools are actually measuring.
- What should a school or employer do instead of relying on a detector score?
- Ask for the artifacts that come from a person's own process — drafts, version history, an ability to explain and extend the reasoning on the spot — rather than a percentage from a tool whose vendors themselves warn against using it as sole evidence.