A peer-reviewed RCT across four universities randomly assigned 3,080 short-answer responses to GPT-4 or human grading. Scores were statistically indistinguishable, and AI let instructors give large classes feedback matching classes a third the size.
-0.04
Difference in grading accuracy, AI vs human (not statistically significant) (2023-2024)
3,080
Student responses randomly assigned to AI or human grading (2023-2024)
~70 students given feedback depth typical of classes a third to a quarter the size
Class size at which AI-assisted feedback matched much smaller classes
Details
Promoter
University of Houston, with University of South Carolina, University of Michigan & Academia Sinica (Taiwan)
Period
2023-2025
Keywords
higher education, political science, AI-assisted grading, research
Context
Grading and providing personalised feedback on short-answer questions is time-consuming, which often pushes instructors of large classes toward multiple-choice assessments and reduces students' opportunities to develop critical-thinking skills.
Objectives
Test whether large-language-model (GPT-4) assistance can grade short-answer responses as accurately as human graders and let instructors give large classes the kind of individualised feedback usually reserved for much smaller classes.
Activities
Researchers led by Heinrich and colleagues ran a randomized controlled trial across four undergraduate political-science programmes (University of Houston, University of South Carolina, University of Michigan, and Academia Sinica in Taiwan) in 2023-2024. Across 26 tests covering 271 students, roughly ten student responses per short-answer question were randomly assigned to human or GPT-4 grading, with the AI guided by instructor-built rubrics and "gold" example answers.
Results
The study, published in PLoS One in August 2025, found no statistically significant difference in grade accuracy between AI and human graders (mean difference -0.04). Regrade requests were only marginally higher for human-graded work, and human feedback was rated "helpful" about two percentage points more often than AI feedback. The clearest gain was productivity: AI-assisted grading let instructors deliver individualised feedback to classes of roughly 70 students at a depth typically reserved for classes a third to a quarter that size.
Conclusions
The authors are explicit about limits: the trial measured grading accuracy and feedback ratings, not actual student learning gains, results varied across the four courses, and reliance on a single proprietary model (GPT-4) raises replicability concerns as models change.
Implementation
Indicative cost
Low (< €50k)
Time to results
Medium (1–3 years)
Staffing & skills
Course instructors building rubrics and gold-standard example answers, Research team (Heinrich et al.) administering random assignment and analysis
Conditions for success
Instructor-built rubrics and gold example answers to guide AI grading
Random assignment of responses across both AI and human graders for comparability
Multi-institution design across differing course contexts
An RCT across four political science courses at Houston and partner universities found GPT-4-assisted grading of short-answer questions, using instructor …