evidoria

← Back to browse

Good practice Imported

Grading at Scale — A Randomized Trial of AI vs. Human Graders Across Four Political-Science Programs

United States of America · Houston · See the United States of America profile

Evidence: Randomised controlled trial Top 17% 67/100 · Ask Evidence Copilot about this practice

A peer-reviewed RCT across four universities randomly assigned 3,080 short-answer responses to GPT-4 or human grading. Scores were statistically indistinguishable, and AI let instructors give large classes feedback matching classes a third the size.

-0.04
Difference in grading accuracy, AI vs human (not statistically significant) (2023-2024)
3,080
Student responses randomly assigned to AI or human grading (2023-2024)
~70 students given feedback depth typical of classes a third to a quarter the size
Class size at which AI-assisted feedback matched much smaller classes
Grading at Scale — A Randomized Trial of AI vs. Human Graders Across Four Political-Science Programs

Details

Promoter
University of Houston, with University of South Carolina, University of Michigan & Academia Sinica (Taiwan)
Period
2023-2025
Keywords
higher education, political science, AI-assisted grading, research

Context

Grading and providing personalised feedback on short-answer questions is time-consuming, which often pushes instructors of large classes toward multiple-choice assessments and reduces students' opportunities to develop critical-thinking skills.

Objectives

Test whether large-language-model (GPT-4) assistance can grade short-answer responses as accurately as human graders and let instructors give large classes the kind of individualised feedback usually reserved for much smaller classes.

Activities

Researchers led by Heinrich and colleagues ran a randomized controlled trial across four undergraduate political-science programmes (University of Houston, University of South Carolina, University of Michigan, and Academia Sinica in Taiwan) in 2023-2024. Across 26 tests covering 271 students, roughly ten student responses per short-answer question were randomly assigned to human or GPT-4 grading, with the AI guided by instructor-built rubrics and "gold" example answers.

Results

The study, published in PLoS One in August 2025, found no statistically significant difference in grade accuracy between AI and human graders (mean difference -0.04). Regrade requests were only marginally higher for human-graded work, and human feedback was rated "helpful" about two percentage points more often than AI feedback. The clearest gain was productivity: AI-assisted grading let instructors deliver individualised feedback to classes of roughly 70 students at a depth typically reserved for classes a third to a quarter that size.

Conclusions

The authors are explicit about limits: the trial measured grading accuracy and feedback ratings, not actual student learning gains, results varied across the four courses, and reliance on a single proprietary model (GPT-4) raises replicability concerns as models change.

Implementation

Indicative cost
Low (< €50k)
Time to results
Medium (1–3 years)
Staffing & skills
Course instructors building rubrics and gold-standard example answers, Research team (Heinrich et al.) administering random assignment and analysis

Conditions for success

  • Instructor-built rubrics and gold example answers to guide AI grading
  • Random assignment of responses across both AI and human graders for comparability
  • Multi-institution design across differing course contexts

Commonly funded by

Horizon Europe Erasmus+ National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful