Which AI detector actually catches Turnitin-level stuff?

Been doing some digging lately because a client of mine runs content for a university-adjacent media outlet, and they’ve started requiring that pieces pass a Turnitin check before publication. Not a full academic submission thing, just a flag for anything that reads AI-generated. I’m used to running checks before delivery anyway, but this one’s got me thinking more carefully about which detector is actually calibrated closest to what Turnitin flags.

My usual workflow involves a couple of tools but I honestly don’t know how well they correlate with Turnitin specifically. A piece can come back clean on one and then get flagged somewhere else entirely. I’ve had that happen twice in the last month.

So I’m curious. Has anyone actually run tests comparing results across detectors vs. what Turnitin actually throws up? I’m not looking for a ranked list, I just want to understand if there’s a pattern. Do certain detectors lean more toward Turnitin’s behavior? Is the logic similar enough that a clean result somewhere means anything?

This feels like something people have experimented with but don’t necessarily write up anywhere useful.

The practical answer for a client workflow is to treat Turnitin as its own separate check and not assume your standard tool covers it. If the client specifically requires Turnitin clearance, that’s the test you need to run, not a proxy for it. Learned that the hard way on a B2B content contract where the client’s legal team ran their own screen. Proxies don’t hold.

Yeah I’ve run this comparison more than I’d like to admit. Short version: the correlation is real but inconsistent. A piece that clears one tool completely can still get flagged by Turnitin, especially if it’s got academic structure, dense citation patterns, or very even sentence length throughout. Turnitin seems to weight consistency of rhythm more than some others do. I’ve started treating a clean result on any one tool as necessary but not sufficient.

The mismatch between tools is documented, it’s just not documented anywhere convenient. Most of the comparison data that exists is either from the companies themselves or from researchers who tested small samples under controlled conditions that don’t reflect real editorial workflows. If you find a consistent pattern that holds across a dozen pieces of mixed content, write it up. Genuinely, there’s not enough of that.

From the teaching side this is actually a real issue. I’ve seen both directions. Students whose work reads as suspicious but passes the tool, and writing that gets flagged that I’m confident is genuine. The calibration difference matters enormously when you’re the one making a judgment call. I don’t use detector output alone anymore. It’s one data point.

we literally had a teacher use turnitin on something and the other tool she ran it through said nothing. same doc. no edits. so yeah it’s not like they’re using the same system at all lol