Tools & Reviews · Posted by Carlos Mendes ·

Comparing GPTZero, Originality.ai and Turnitin on the same batch of essays

0

ran an experiment over the last two weeks of summer school. took forty student essays, half written entirely by students and half generated fully by chatgpt with light editing, then ran all forty through gptzero, originality.ai and turnitin’s ai writing indicator. fascinating results honestly. originality.ai flagged the ai generated batch most accurately at around 85 percent, gptzero was close behind but had more false positives on the human written essays, especially ones with a formal tone. turnitin’s ai indicator undercounted the ai batch significantly, missing almost a third of the fully generated essays. none of the three were reliable enough to use as a standalone verdict but the variance between them surprised me more than i expected going in. happy to share the raw spreadsheet if anyone wants to replicate this with their own students work.

5 replies

5 Replies

0

forty essays is a decent start but i'd want to see this replicated with a bigger sample before drawing conclusions about which tool is actually more accurate. did you control for essay length and subject area? formal tone essays skewing gptzero false positives could just as easily be a length effect.

0

@Daniel Kowalski fair pushback, length wasnt controlled for as tightly as i'd like, essays ranged from 400 to 900 words. planning a bigger run this term with matched lengths across both groups, will post an update once i have it.

0

my own writing got flagged at 38% by gptzero once when i ran a sample paragraph through just to test it, and i've been teaching english for eleven years. these tools have a real problem with anyone who writes cleanly and formally, human or not.

0

honestly i doubt any of these tools will ever get reliable enough for high stakes decisions. the underlying technology is chasing a moving target since the writing models keep improving too. be careful leaning on any of the three for actual discipline.

0

still really useful data though even with the caveats. bookmarking this for our department meeting next week, we've been arguing about which tool to pilot for months.