Are AI detectors reliable?
Not reliable enough to be treated as evidence. A 2023 study of fourteen tools found none reached 80 percent overall accuracy, and accuracy collapsed on text that had been edited or paraphrased. A detector score is an investigative signal, not proof that anything was generated.
Updated
The published evidence is consistent and unflattering. Weber-Wulff and colleagues tested fourteen detection tools, including widely used commercial ones, against human text, raw generated text, human-edited generated text, and machine-paraphrased text. No tool reached 80 percent overall accuracy and only five exceeded 70 percent. Performance on human text was good, around 96 percent, and it fell to roughly 42 percent on edited AI text and 26 percent on paraphrased AI text.
That pattern is the important finding. Detectors are best at catching text nobody bothered to edit, which is the least likely case in academic work, and worst at catching text that was revised, which is the most likely. Separate research on paraphrasing attacks reported detection of watermarked text falling from over 99 percent to about 4 percent at a fixed false-positive rate.
Detectors do perform well under narrow, controlled conditions, which is why the vendor numbers are not invented. Gao and colleagues generated 50 abstracts from the titles of real papers in five high-impact medical journals, and the GPT-2 Output Detector gave those generated abstracts a median fake score of 99.98 percent against 0.02 percent for the originals, with an AUROC of 0.94 (Gao et al., 2023, p. 1). That is one prompt style, one genre, and unedited output, none of which describes a thesis chapter you wrote and revised over three months.
The same study measured people against the tool on the same texts. Four blinded reviewers correctly identified 68 percent of the generated abstracts and 86 percent of the real ones, so they missed about a third of the generated abstracts and wrongly labeled 14 percent of the authentic ones as generated (Gao et al., 2023, p. 2). Detector scores did not differ significantly between generated abstracts the reviewers caught and generated abstracts they judged authentic, at p = 0.45, which means the tool and the reader were responding to different cues rather than confirming each other (Gao et al., 2023, p. 3).
Vendors report better numbers on their own test sets, and those numbers are not fabricated, they are measured under a chosen corpus and threshold. What they do not establish is the false-positive rate for a particular student, in a particular discipline, writing a particular kind of text.
The conclusion most institutions have reached is that a score can start a conversation and cannot end one. Authorship questions are settled by drafts, notes, version history, and the ability to explain your own work.
Why do detectors flag non-native English speakers more often?
Because they key on low linguistic variability, which is also a feature of second-language writing. In a 2023 study in Patterns, seven detectors misclassified 61 percent of human-written TOEFL essays as AI-generated on average, while classifying native-speaker essays almost perfectly.
In the same study, all seven detectors unanimously misclassified about 20 percent of those essays, and at least one detector flagged nearly 98 percent of them. That is a disparate impact large enough to make detector use on multilingual cohorts a fairness problem rather than a technical one.
The mechanism is not mysterious. Simpler vocabulary and more predictable sentence structure look statistically like generated text, and second-language academic writing frequently has both.
Liang and colleagues tested that mechanism directly rather than assuming it. Rewriting the TOEFL essays with richer vocabulary, so they read more like a native speaker, reduced misclassification, and simplifying the vocabulary of the US student essays increased it, across all seven detectors (Liang et al., 2023, p. 2). The detectors were tracking linguistic proficiency, not authorship. The same authors point out that a vendor claim of 99 percent accuracy cannot be checked by anyone outside the company, because the test datasets, model details, and training data are not published (Liang et al., 2023, p. 2).
What do I do if my own writing gets flagged?
Produce the evidence that a detector cannot: drafts with version history, notes, search logs, an outline, and your ability to explain the argument in conversation. Ask what the score is being used for, since most institutions state that a detector result alone is not sufficient grounds for a finding.
Version history is the strongest single artifact. A document with hundreds of incremental revisions over weeks is very difficult to fabricate and very easy to show, which is a good reason to draft in something that keeps history rather than in a single overwritten file.
Ask, politely, for the false-positive rate the institution is assuming and the evidence for it. That question is legitimate, it is answerable, and in most cases the answer will make clear that the score cannot carry the weight being placed on it.
You also have a published recommendation to cite. Liang and colleagues advise against using GPT detectors in educational or evaluative decisions at all, especially where non-native English writers are involved, and say that any deployment should first be tested on samples from the actual domain and reviewed with domain experts and the people affected (Liang et al., 2023, p. 3). They suggest formative uses instead, such as flagging clichés or repetitive phrasing for revision. That is a reasonable thing to put in writing to whoever is making the decision about your work.
Sources
- Weber-Wulff et al., Testing of detection tools for AI-generated text, International Journal for Educational Integrity (2023) · checked 6 August 2026
- Stanford HAI on Liang et al., GPT detectors are biased against non-native English writers, Patterns (2023) · checked 6 August 2026
- Sadasivan et al., Can AI-generated text be reliably detected?, arXiv (2023) · checked 6 August 2026
- Otterbacher, Why technical solutions for detecting AI-generated content in research and education are insufficient, Patterns (2023) · checked 6 August 2026
- Liang et al., GPT detectors are biased against non-native English writers, Patterns (2023) · checked 6 August 2026
- Gao et al., Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers, npj Digital Medicine (2023) · checked 6 August 2026