Best AI content detector for academic writing. My testing notes

Running a comparison over the last six weeks for a piece I’m writing on AI tools in graduate education. Tested five detectors on the same set of documents: a mix of confirmed AI-generated text, human-written academic prose, AI-assisted text with heavy editing, and my own writing.

Key findings worth sharing. The tools diverge most sharply on the AI-assisted-with-heavy-editing category. Two tools called these mostly human. Two called them mostly AI. One produced inconsistent results on the same document across multiple runs, which tells you something about reliability.

On confirmed human academic writing: false positive rates were higher than I expected across all tools. Formal academic prose, with its passive constructions, citation density, and hedged language, reads as AI to multiple detectors at elevated rates. That’s a real validity problem.

On confirmed AI output with no editing: high agreement across tools, as expected. The clear cases aren’t where these tools differentiate.

The practical implication is that the value of these tools is narrowest exactly where you’d want it to be broadest: in the ambiguous middle. Happy to share more specific notes if useful.

The ‘ambiguous middle’ problem you’re describing is structurally predictable. The cases where detection would actually matter, the thoughtfully AI-assisted work, are exactly where the tools have the least signal. The tools work best on the cases that are easiest to identify anyway.

Please share more notes. This kind of empirical comparison is exactly what’s missing from the public conversation. Most of what’s out there is either vendor marketing or anecdote. Systematic testing with documented methodology is genuinely useful.

The inconsistency across multiple runs on the same document is a significant finding. A reliable classifier should produce stable output on identical input. Variable results suggest the model has high uncertainty in that region of the feature space, which means the output is not interpretable as a meaningful score. Worth highlighting prominently.

The false positive rate on formal academic prose is the finding that matters most for practitioners. If the tools are systematically flagging well-written human academic text, the entire use case for academic integrity falls apart. You’d be investigating the strongest writers. That’s backwards.

From a decision-making standpoint, a tool with inconsistent results on the same input isn’t providing information, it’s providing noise. Any workflow that treats that output as actionable is making consequential decisions on noise. That’s a liability, not a solution.