Do AI Detectors Actually Work? What the Research Shows
OpenAI withdrew its own detector for low accuracy. A study in Patterns found seven commercial detectors misclassified more than half of non-native English essays. Here is what the evidence actually supports.
An AI detector never tells you who wrote a document. It tells you how statistically predictable the text is, then converts that measurement into a percentage that looks far more authoritative than it is. That gap between what the tool measures and what people believe it proves is where most of the harm happens: students accused without evidence, freelancers losing clients, non-native English writers flagged again and again for writing carefully.
This article covers how detectors actually work, why false positives are structural rather than accidental, what a score can and cannot support, and what to do if you are wrongly accused.
How AI detectors work
Almost every text-based detector rests on the same underlying idea: a language model can estimate how likely each word in a sentence is, given the words before it. Text that consistently chooses the most probable next word looks machine-generated. Text that takes unexpected turns looks human.
Two measurements do most of the work.
Perplexity
Perplexity is a measure of how surprised a language model is by a piece of text. Low perplexity means the model would have predicted roughly those words itself. High perplexity means the text kept going somewhere the model did not expect.
Large language models are trained to produce high-probability continuations, so their output tends to sit in the low-perplexity range. That is the signal detectors look for. The problem is that plenty of human writing is also low-perplexity: formulaic academic prose, legal boilerplate, technical documentation, standardised exam essays, and the writing of anyone who has learned English through structured instruction rather than immersion.
Burstiness
Burstiness measures variation. Human writers tend to produce uneven text: a twenty-eight word sentence followed by four words, a dense paragraph followed by a thin one, an abrupt aside. Model output is typically smoother, with sentence lengths and structures clustering around a mean.
So a detector is essentially asking two questions: is this text unusually predictable, and is it unusually even? A high score means yes to both. It does not mean a model produced it.
Classifier-based approaches
Some detectors add a trained classifier on top: feed it large sets of known human text and known AI text, and let it learn whatever features separate them. This works well on the data it was trained on and degrades on anything else, including newer models, edited output, and writing from populations underrepresented in the training set. It is also opaque. When a classifier flags a paragraph, there is usually no explanation beyond the number.
The accuracy problem is documented, not hypothetical
Two well-known data points are worth stating precisely, because they are frequently misquoted.
OpenAI withdrew its own detector. In January 2023 OpenAI released an "AI Text Classifier" intended to distinguish AI-written from human-written text. In July 2023 the company discontinued it, citing its low rate of accuracy. The organisation with the most direct knowledge of how its models generate text could not build a reliable detector for that text and said so publicly.
Detectors are biased against non-native English writers. A study by Liang and colleagues, published in the journal Patterns in 2023, tested several widely used GPT detectors on TOEFL essays written by non-native English speakers and on essays written by native English speakers. The detectors classified the native-speaker essays accurately but misclassified a majority of the non-native writers' essays as AI-generated. The researchers linked this to the same statistical properties described above: writing that uses a narrower vocabulary range and more conventional sentence construction reads as low-perplexity, whether a person or a model produced it.
That finding matters enormously in Bangladesh, India, Pakistan, Nigeria, the Philippines and every other context where excellent English writers learned the language formally. The penalty falls hardest on people who write carefully and conventionally, which is precisely what academic instruction teaches.
Turnitin suppresses low scores. Turnitin's AI writing indicator does not display results below 20 percent, because the company has acknowledged that false-positive rates are higher in that band. This is a sensible design decision, and it is also an admission: the tool's own vendor treats its low-range output as unreliable enough to hide.
Why false positives are structural
It is tempting to assume detectors will improve until the error rate approaches zero. There are reasons to doubt that.
- The target keeps moving. Detectors are trained on the output of models that already exist. Each new model generation shifts the statistical profile, and detectors lag behind.
- The signal overlaps with legitimate writing. Low perplexity is not an artifact of machine generation. It is a property of conventional, clear, well-structured prose. There is no way to flag one without flagging the other.
- Editing collapses the signal. Substantive human revision of an AI draft changes exactly the properties detectors measure. So does an AI draft that a person has rewritten heavily. So does a human draft polished with grammar software. The categories detectors assume are clean are, in practice, a spectrum.
- Base rates make errors worse than they look. Even a detector with a genuinely low false-positive rate produces a large absolute number of wrong accusations when run across thousands of submissions. A one percent error rate across ten thousand essays is one hundred students accused wrongly.
What a detector score does and does not prove
| A high score suggests | A high score does not prove |
|---|---|
| The text is statistically predictable and rhythmically even | That an AI model generated it |
| The prose may be worth a closer human read | That the named author did not write it |
| The writing follows conventional patterns | That academic misconduct occurred |
| A conversation with the writer may be useful | Anything at all about intent |
The correct way to read a score is as a prompt for human judgement, never as a verdict. A detector output is closer to a smoke alarm than to a fingerprint: worth investigating, useless as sole evidence.
A probability is not a finding. A tool that outputs "92% AI" has not identified an author. It has reported that the text sits in a statistical region where model output often sits, and where a great deal of human writing also sits.
What to do if you are wrongly accused
If a detector has flagged work you wrote, the goal is to shift the conversation from a number to evidence of process. Detector scores are hard to argue with directly. Drafting history is not.
- Do not admit to something you did not do. People under pressure sometimes concede partial responsibility to end an uncomfortable meeting. That concession becomes the record.
- Produce your version history. Google Docs keeps revision history. Microsoft Word tracks versions in OneDrive and SharePoint. Overleaf, Notion and most modern editors keep an edit timeline. A document that grew over eleven sessions across two weeks is strong evidence of authorship.
- Gather your working materials. Handwritten notes, annotated PDFs, browser history for sources, outline drafts, messages to a supervisor asking about scope. These form a chain that no detector output can contradict.
- Ask what the score means in the institution's own policy. Many universities have written that AI-detection output cannot alone constitute evidence of misconduct. Ask for the policy in writing and read what it commits the institution to.
- Ask about the false-positive rate. Ask which tool was used, what threshold triggered the accusation, and what the vendor publishes about accuracy for writers whose first language is not English. The Liang study is a reasonable thing to cite.
- Offer a viva. Volunteering to discuss the work in person is persuasive because someone who wrote a piece can explain the choices behind it. Someone who did not, usually cannot.
- Get a second reader. A student union representative, an academic adviser or a colleague changes the dynamic of the meeting and creates a witness.
Why "beating the detector" is the wrong frame
A large industry now sells tools that promise to make text undetectable. Most of them work by paraphrasing: swapping synonyms, restructuring clauses, adding irregularity. This raises perplexity, which lowers the score, and often makes the writing worse. Awkward synonym choices, broken idioms and mangled technical terms are common outputs.
It also does nothing about the underlying issue. If the concern is that a piece of writing carries no thinking, a paraphrase carries no thinking either. If you are weighing this question, it is worth understanding what bypassing AI detection actually involves before assuming a tool solves the problem.
The version of this that holds up is editing rather than obfuscation: adding what the draft lacks, cutting what it padded, checking the claims, and rewriting in a voice that is yours. That changes the detector score as a side effect, because genuinely revised text has different statistical properties. But the reason to do it is that the writing gets better.
Where this leaves things
AI detectors measure something real. Model output does tend to be more predictable and more even than human writing, on average. The measurement is not fake.
What is false is the leap from measurement to accusation. No current detector can establish who wrote a document. The vendor of the most widely deployed tool hides its own low-range scores. The company that builds the most widely used language models withdrew its detector for poor accuracy. Peer-reviewed work has shown the error falls disproportionately on non-native English writers, which makes uncritical use of these tools a fairness problem as well as an accuracy problem.
If you run an institution, treat scores as a trigger for a conversation and never as evidence on their own. If you are a writer, keep your drafting history: it is the only thing that reliably settles the question. And if you are editing AI-assisted work, judge the result by whether it says something worth reading, not by what a percentage bar reports.
More from the blog
Using AI for a Literature Review Without Wrecking It
AI has a real role in a literature review, and it sits later than most people put it: afte...
Flagged as AI for Writing in Your Second Language: What to Do
It is not your imagination. Detectors score vocabulary range, not authorship, and a peer-r...
Does Google Penalise AI Content? What the Policies Actually Say
No, Google does not penalise AI-written content. It penalises pages published at scale to...