I work with writers professionally and I’ve been thinking about who actually gets hurt by false positives in AI detection.
The writers most at risk aren’t the ones using AI heavily. They’re the ones who write cleanly. Precise syntax. Efficient structure. Clear transitions. Consistent register. These are the hallmarks of good craft and they’re also, apparently, the features that make detectors nervous.
The writers who sail through detection tend to have more casual, uneven prose. Typos, sentence fragments, tonal inconsistency. The detector reads these as human tells. In effect, the tools are penalizing writing quality.
I’m not making this point to argue that detectors are useless. I’m making it because I think the people designing workflows around detection need to understand this dynamic. The writers you most need to trust, the polished professionals, are the ones most likely to get flagged. The writers you should be scrutinizing more carefully are the ones who know how to write badly on purpose.
I’ve stopped using detector scores for any writer whose baseline I already know is strong. The false positive rate is too high and the consequences of a wrong accusation are too severe. The tool is only useful for students whose baseline I don’t have, and even then only as one signal among several.
this is kind of how a lot of automated systems work though. the people who know how the system works game it. the people who don’t know get caught even when they’re doing the right thing. it’s not specific to AI detection
The ‘writing badly on purpose’ angle is one I think about from an SEO content perspective. Writers who understand detectors can produce text that reads as human while using AI heavily. Writers who don’t understand detectors get flagged for writing well. The tool punishes the unaware more than the sophisticated.
This is something I’ve observed directly in manuscript review. The clean, precise writing we’d normally want to acquire flags at higher rates than rougher, more idiosyncratic prose. It creates a perverse incentive at the point of detection: the better the writing, the more suspicious it looks.
The calibration bias you’re describing is documented in the academic literature on detector evaluation. Tools trained predominantly on clearly AI-generated versus clearly human text have a blind spot in the middle. Polished, formal, consistent human writing sits in that blind spot. It’s a structural problem with the training approach.