AI Detection · Posted by Tanya Whitfield ·

Turnitin AI Detection – How Has It Actually Changed Over 3 Years?

6

I’ve been using Turnitin since 2008. I remember when the plagiarism detection was inconsistent and unreliable and how it gradually improved. I’ve been watching the AI detection feature with the same patience and trying to assess whether it’s following the same trajectory.

Honest assessment after 3 years of watching:

2023 (AI detection launch): very high false positive rate. Turnitin themselves acknowledged limitations. Useful for research, not for formal decisions.

2024: false positive rate improved somewhat. Still unacceptable for high-stakes use. Turnitin added more nuanced scoring.

2025-2026: meaningful improvement on non-humanized AI text. Still problematic on ESL writing. Humanized text largely undetected. The arms race is clearly slowing progress.

The trajectory is improving but slower than plagiarism detection improved, because humanizers actively work to defeat it.

My prediction from 2023 was that it would take 5 years to reach the reliability level that plagiarism detection reached. I still think that’s roughly right – 2027-2028 before it’s reliable enough for formal integrity decisions. We’re not there yet.

3 replies

3 Replies

9

The historical framing is useful for calibrating expectations. Plagiarism detection accuracy also took years. The difference: there was no active plagiarism-defeating technology evolving simultaneously. The AI/humanizer arms race is a structural difference that may permanently cap detection accuracy below what plagiarism detection achieved.

4

the 2027-2028 estimate assumes humanizer technology doesn't keep pace. if it does, Turnitin's accuracy ceiling may be lower than plagiarism detection achieved. different adversarial dynamics.

8

Turnitin's AI detection has gotten more aggressive over the past year in my experience. more flagging, not necessarily more accurate flagging. the pressure to show detection capability seems to have shifted the threshold. it's worth calibrating your own baseline before using it for consequential decisions - what it flagged at 50% probability a year ago is now flagging at 70% on similar content.