evidoria

← Back to browse

Good practice Imported

AI-Assisted Grading and Personalized Feedback in Large Political Science Classes — A Multi-University Randomized Trial from Houston

United States of America · Houston · See the United States of America profile

Top 43% 53/100 · Ask Evidence Copilot about this practice

An RCT across four political science courses at Houston and partner universities found GPT-4-assisted grading of short-answer questions, using instructor rubrics, matched human grading quality without raising regrade requests, easing pressure toward multiple-choice-only tests.

Details

Promoter
University of Houston
Period
2023-2024 data collection; published August 2025
Keywords
higher education, generative AI, assessment, political science

Description

Grading short-answer questions at scale is one of the most time-consuming parts of teaching large university courses, which pushes many instructors toward multiple-choice tests that give students less room to develop and demonstrate critical thinking. A team from the University of Houston, University of South Carolina, Academia Sinica, the University of Michigan and Kangwon National University tested whether GPT-4 could close that gap without sacrificing grading quality.

Across four large political science courses (271 students, roughly 3,080 short-answer responses, 88% of them AI-graded), instructors built prompt templates containing their rubric, worked examples with feedback, and each new student response. The team then benchmarked the AI-assisted grades against instructor grading on measures including how well scores tracked a student's performance elsewhere in the course, how often students requested regrades (a proxy for grading errors), and how helpful students found the feedback.

The randomized trial found AI-assisted grading approximated what an instructor grading a small class would produce, without an increase in regrade requests, while freeing instructor time. The authors are explicit that this is a productivity and consistency tool for feedback delivery, not a replacement for instructor judgment - human review remained part of every grading pipeline tested.

Read the full analysis: https://sanghoon-park.com/publications/articles/heinrich-et-al-2025/

Implementation

Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful