AI detectors are not accurate enough to base a cheating accusation on, and the research backs that up: some tools miss a large share of AI-written text outright, and several flag non-native English writers at rates several times higher than native speakers. If a detector is the only thing standing between a student and a zero, that is a problem before it is a solution.
Mr. Ortiz found this out the hard way last spring. A detector flagged a student's reading response as "98 percent AI-generated." The student, an English learner who had rewritten her ideas twice to get the wording right, was furious, and rightly so. The essay was hers. The detector was wrong. Mr. Ortiz spent the next two class periods rebuilding trust he never should have had to lose.
Why AI detectors keep getting it wrong
Detectors work by guessing at statistical patterns, things like sentence predictability and word choice, that tend to show up in machine-generated text. The trouble is that careful, methodical writers, especially English learners who were taught formulaic structures, produce those same patterns naturally. As Edutopia has reported, a 2026 follow-up to the original Stanford research found a mean false positive rate of 61 percent for TOEFL essays from Chinese students, compared with about 5 percent for essays from US students in the same test. That is not a rounding error. That is a tool that is systematically wrong about the students who can least afford to be wrongly accused.
Detectors also struggle with anything that is not pure, unedited AI output, which describes most real cheating. A student who drafts with a chatbot and then edits by hand produces "hybrid" text, and accuracy on hybrid writing drops close to zero in independent testing. The students most worth catching, in other words, are the ones detectors are worst at catching. It is one reason researchers writing in The Conversation argue that detection software was never going to be the fix, and that assessment itself has to shift.
What actually happens when a student outsources the thinking
The real damage of AI-assisted cheating is not that a student turned in text they did not type. It is that they turned in an idea they never actually held in their head. A finished paragraph about photosynthesis proves nothing about whether the student could explain photosynthesis to a classmate, or notice when their own explanation contradicts itself. That gap, between a polished artifact and actual understanding, is what a detector was never built to measure in the first place. This is the same gap covered in our earlier look at students outsourcing thinking to AI tools, and it is worth revisiting here because it is the real problem, not the writing itself.
This is why assessment design, not detection software, is where the real leverage sits. A prompt that only asks for a finished answer rewards anyone who can produce a finished answer, by any means. A prompt that asks a student to show their thinking, revise it, or explain it out loud narrows the gap considerably. It does not close it, exactly, but it makes shortcutting a lot more work than just doing the assignment honestly.
Four questions that make an assignment harder to fake
Before assigning anything, it helps to run it through a short check. None of these questions require new software, just a different prompt.
- Does it require a specific source? A prompt tied to a passage students actually read, with a quote or detail only that passage contains, is much harder to shortcut than a generic essay question.
- Is there a visible draft or process step? Asking for an annotated outline, a voice memo, or a marked-up first draft before the final version means there is more than one artifact to fake.
- Can the student defend it out loud, on the spot? A two-minute follow-up conversation, even an informal one, tends to reveal in seconds whether a student understands what they submitted.
- Would a stranger's answer sound the same? If the prompt could be answered identically by any student anywhere, it is also answerable identically by a chatbot.
None of these fixes are dramatic on their own. Stacked together across a unit, though, they change the math for a student who is deciding whether shortcutting is even worth the effort. For more on what a rubric-level version of this looks like, see our notes on oral assessment best practices.
What this looks like in an oral assessment
The "defend it out loud" question above is worth taking seriously, because it is the one detectors cannot touch at all. When ArticulAI runs a short adaptive oral check-in after a written assignment, it is not re-grading the essay. It is asking a follow-up question the essay did not answer, then another one based on what the student says back. A student who wrote their own paragraph on photosynthesis can usually talk through it, even messily. A student who pasted in a paragraph they do not understand runs out of things to say within the first follow-up, and that gap tends to show up faster in a thirty-second exchange than it ever would in an essay's word choice.
This will not replace judgment, and it should not try to. A teacher who knows their students still catches things a transcript never will. But a follow-up question is a much cheaper, faster check than reading every essay for suspicious phrasing, and it does not carry the same risk of falsely accusing a student whose only "crime" was writing carefully.
None of this requires abandoning writing assignments or treating every student like a suspect. It just means shifting some of the weight off a piece of software that was never built to carry it, and back onto the kind of quick check a good teacher already does instinctively: asking a student to explain what they meant. That's the exact moment ArticulAI is built for. See how ArticulAI works.
Frequently asked questions
Are AI detectors accurate enough to catch cheating?
Not reliably. Independent testing in 2026 found accuracy ranging from roughly 65 to 90 percent depending on the tool, with performance on edited or hybrid AI text dropping close to zero. No major detector has cleared 80 percent accuracy across all writing types.
How can a teacher tell if a student used ChatGPT without running a detector?
Ask the student to explain or defend the work out loud. A short follow-up question about a specific claim in the writing usually surfaces the gap between what was submitted and what the student actually understands.
What makes an assignment AI-resistant?
Assignments tied to a specific source, a visible draft step, or a spoken defense are harder to shortcut than a generic prompt, because there is more than one artifact for a student to fake and more chances for gaps in understanding to show.
Do AI detectors falsely flag non-native English speakers?
Yes, disproportionately. Research has found false positive rates for non-native English writers several times higher than for native speakers, largely because formulaic, careful writing patterns overlap with what detectors associate with AI-generated text.
What should a teacher do if a detector flags a student's work?
Treat the flag as a starting point for a conversation, not a verdict. Ask the student to walk through their process or explain a section out loud before assuming the flag is accurate, since false positives are common enough to warrant that step every time.

