evidoria

← Back to browse

Good practice

Vanderbilt University's Evidence-Based Rejection of Turnitin's AI-Writing Detector

United States of America · Nashville · See the United States of America profile

Vanderbilt disabled Turnitin's AI-writing detector campus-wide in Aug 2023, calculating its claimed 1% false-positive rate could wrongly flag ~750 of its 75,000 annual papers; a Stanford study found similar detectors misjudged 61.3% of non-native TOEFL essays as AI-written.

1 %
Turnitin's claimed false-positive rate (2023)
75000 papers
Papers submitted through Turnitin annually (2022)
750 papers/year
Estimated papers wrongly flagged per year at claimed false-positive rate (2022)
61.3 %
Average misclassification rate of non-native TOEFL essays as AI-written across seven detectors (2023)
97.8 %
Highest single-detector misclassification rate of non-native TOEFL essays (2023)
19.8 %
TOEFL essays unanimously misjudged as AI-written by all seven detectors (2023)
11.6 %
Misclassification rate after using ChatGPT to enhance non-native essays' word choice (2023)
91 essays
TOEFL essays analysed in the Stanford study (2023)
Vanderbilt University's Evidence-Based Rejection of Turnitin's AI-Writing Detector

Details

Promoter
Vanderbilt University
Period
Apr 2023 – Aug 2023 (decision); bias research published Jul 2023
Keywords
higher education, academic integrity, AI-text detection, plagiarism, writing assessment, EdTech governance

Context

In April 2023 Vanderbilt University's Brightspace/Center for Teaching enabled Turnitin's newly launched AI-writing detection feature for faculty use campus-wide, even though Turnitin publicly claimed only a roughly 1% false-positive rate. With about 75,000 papers submitted through Turnitin in 2022, Vanderbilt staff calculated that even this low claimed rate implied around 750 papers a year could be wrongly flagged as AI-generated. That same summer, a Stanford University research team published a peer-reviewed study in the Cell Press journal Patterns testing whether commercial AI-text detectors were biased against non-native English writers.

Objectives

Vanderbilt's stated aim was to assess whether Turnitin's AI-writing detector was accurate and fair enough to underpin high-stakes academic-integrity decisions, weighing the tool's claimed false-positive rate against the number of papers it screened and the transparency of its methodology.

Activities

Over several months, Vanderbilt's Brightspace/Center for Teaching ran internal testing of the AI-writing detector and consulted with other universities and with AI vendors, including Turnitin, before deciding on its future use. In parallel, Stanford researchers (Liang, Yuksekgonul, Mao, Wu and Zou) ran 91 genuine TOEFL essays written by non-native English speakers, and a comparison set of US eighth-grade essays by native speakers, through seven widely used commercial GPT detectors, and additionally tested how ChatGPT-based rewording of the essays changed detector outcomes.

Results

Vanderbilt disabled Turnitin's AI-writing detector campus-wide on 16 August 2023, citing both the false-positive risk and Turnitin's refusal to disclose the detailed methodology behind its AI-probability scores. The Stanford study found the seven detectors misclassified an average of 61.3% of non-native TOEFL essays as AI-written, with one detector flagging 97.8% and all seven unanimously misjudging 19.8% of the set, while native-English eighth-grade essays were correctly classified; using ChatGPT to enhance the non-native essays' word choice reduced the false-positive rate to 11.6%, whereas simplifying the native essays' vocabulary increased their misclassification.

Conclusions

Together the two cases show a rare feedback loop between an institution's internal cost-benefit reasoning and independent academic research reaching the same conclusion: commercial AI-text detectors were not yet reliable or fair enough for high-stakes academic-integrity decisions, particularly for international and non-native-English-speaking students, prompting Vanderbilt and other US universities to disable AI detection rather than risk false accusations.

Implementation

Indicative cost
Low (< €50k)
Time to results
Short (< 1 year)
Staffing & skills
Vanderbilt's Brightspace / Center for Teaching team, which administered the campus-wide Turnitin AI-detector rollout and later communicated the decision to disable it, Michael Coley, the Vanderbilt staff member who publicly summarised the institutional reasoning for disabling the detector

Conditions for success

  • Willingness to independently quantify a vendor's claimed false-positive rate against the institution's actual submission volume before relying on the tool
  • Consultation with peer universities and with AI-detection vendors, including Turnitin, before making the final decision
  • Attention to independent peer-reviewed bias research (the Stanford Patterns study) alongside the university's own internal calculations
  • Willingness to discontinue a widely adopted EdTech tool when its vendor would not disclose the methodology behind its scores

Common failure modes

  • Relying on a vendor's marketed accuracy figures without checking them against actual submission volume
  • Continuing to use black-box AI-detection tools whose vendors refuse to disclose their underlying methodology

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful