AI detectors on student essays. What are teachers actually seeing?

I’ve been using an ai detector for student essays for a full semester now and I want to be honest about what the experience has actually been, because I think the public conversation around this is still mostly theoretical.

The tool catches obvious cases. The completely unedited GPT output, the weirdly formal five-paragraph essays that don’t match anything the student has produced before. That part works. You get a flag and it more or less aligns with what your gut was already telling you.

The problem is everything in between. I have students who write formally because that’s how they were taught. ESL students whose grammar is unusually clean. Students who outline carefully and write in a structured way. All of these trigger elevated scores. And on the other end, I have students who clearly used AI but ran it through something that cleared the detector.

So I’m using the tool, but I’m also aware that what it’s doing is validating my prior suspicions rather than actually surfacing cases I’d have missed. Whether that’s worth the time investment is a question I’m still sitting with.

What are others actually seeing after a full term of using these tools?

from a student side, the inconsistency is genuinely confusing. like some teachers are scanning everything and others aren’t running anything. there’s no clarity on what the standard actually is. makes it hard to know what you’re supposed to do

The ESL false positive problem is the one that keeps me up as a TA. I’m grading undergrads whose first language isn’t English, and the tools flag them at higher rates than native speakers. If I acted on every flag I’d be disproportionately investigating the students who are already navigating the most barriers. I’ve basically stopped using scores as triggers and only use them as one input among several.

This matches almost exactly what I found. The detector is good at catching the ones you already suspected. It adds very little for the cases in the middle. The false positive rate on ESL students is a consistent and documented problem and I’ve started manually excluding international student submissions from any flagging workflow as a result.

The pattern you’re describing, the tool confirming suspicion rather than generating new signal, is a known limitation in the detection literature. It’s called confirmation bias in automated systems. The tool isn’t wrong, it’s just not doing what most users assume it’s doing. It amplifies existing judgment rather than replacing it. That’s fine if you know it, dangerous if you don’t.

The confirmation bias point is really important. Any tool that works by surfacing what you already suspected isn’t giving you new information, it’s giving you a paper trail. For some institutional contexts that’s the actual value. For actually understanding student writing behavior it’s close to useless.