I’ve been running an informal test over the past two months because I couldn’t justify paying for a detection tool out of pocket and wanted to know whether the free options were worth bothering with at all. Not a rigorous study, just me, three free detectors, and a set of student essays I’d already reviewed manually.
Here’s what I found. All three tools flagged the same two essays that I was already suspicious of. So far so good. But they disagreed on four other submissions where I had no particular concern. One tool called them mostly AI, another said mostly human, third was somewhere in the middle. Same documents.
The false positive problem is real and specific. The students whose work got flagged incorrectly tended to be my stronger writers. Very clean syntax, tight structure, efficient word choice. The tools seem to punish exactly the things good writing coaching teaches.
I’ve stopped using these results as anything other than a prompt to look more closely. If a tool flags something, I read the essay again and compare it to that student’s in-class writing. That comparison is more useful than any score. But I wanted to put the data out there because I think a lot of teachers are assuming these tools are more reliable than they are.
The false positive pattern you describe is consistent with what’s in the literature. Non-native English writers and writers with formal academic training are disproportionately flagged. The tools were not calibrated against diverse writing populations and the bias shows. This is not a minor calibration issue. It’s a structural validity problem.
Two months of informal testing is still more methodologically honest than most of what I’ve seen published by the tool vendors themselves. Appreciate the write-up.
I’m not in education but this is interesting from a trust standpoint. If I were a parent and found out my kid’s work was flagged by a tool that had this level of disagreement across platforms, I’d want to know what the school’s review process was before anything happened. It sounds like you have a thoughtful one, which is reassuring.
This matches my experience almost exactly. The tools do catch the obvious cases. The problem is they also flag competent writing, which creates a due process issue. Any institution relying on detector scores as standalone evidence is making a policy mistake.
The in-class comparison method you mention is what I’ve heard called a ‘triangulation approach’ in some of the academic integrity policy discussions. It’s basically the right move. No single data point is sufficient. A detector flag plus unusual deviation from prior work plus inability to discuss the content is a credible case. Any one of those alone is not.