Gradescope — AI-assisted grading and feedback for higher education
United States of America
AI platform grouping similar student responses for batch grading; 3.2M students at 2,600+ universities. A peer-reviewed ACM 2017 study found …
United States of America · Houston · See the United States of America profile
Top 43% 53/100 · Ask Evidence Copilot about this practice
An RCT across four political science courses at Houston and partner universities found GPT-4-assisted grading of short-answer questions, using instructor rubrics, matched human grading quality without raising regrade requests, easing pressure toward multiple-choice-only tests.
Grading short-answer questions at scale is one of the most time-consuming parts of teaching large university courses, which pushes many instructors toward multiple-choice tests that give students less room to develop and demonstrate critical thinking. A team from the University of Houston, University of South Carolina, Academia Sinica, the University of Michigan and Kangwon National University tested whether GPT-4 could close that gap without sacrificing grading quality.
Across four large political science courses (271 students, roughly 3,080 short-answer responses, 88% of them AI-graded), instructors built prompt templates containing their rubric, worked examples with feedback, and each new student response. The team then benchmarked the AI-assisted grades against instructor grading on measures including how well scores tracked a student's performance elsewhere in the course, how often students requested regrades (a proxy for grading errors), and how helpful students found the feedback.
The randomized trial found AI-assisted grading approximated what an instructor grading a small class would produce, without an increase in regrade requests, while freeing instructor time. The authors are explicit that this is a productivity and consistency tool for feedback delivery, not a replacement for instructor judgment - human review remained part of every grading pipeline tested.
Read the full analysis: https://sanghoon-park.com/publications/articles/heinrich-et-al-2025/
Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.
Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.
Where this practice's information was retrieved from, and when.
United States of America
AI platform grouping similar student responses for batch grading; 3.2M students at 2,600+ universities. A peer-reviewed ACM 2017 study found …
United States of America
A peer-reviewed RCT across four universities randomly assigned 3,080 short-answer responses to GPT-4 or human grading. Scores were statistically indistinguishable, …
Mexico
Tec de Monterrey AI ecosystem (33 campuses, 2023): CogBooks adaptive learning (Springer 2024, +15% A-grades), Adaptive Leveling Modules (10,350 students, …
South Africa
Pretoria accounting lecturers built a GPT-4 web tool giving essay-question feedback engineered around Nicol & Macfarlane-Dick's principles. A study of …
Open full copilot Grounded in cited practices — always check the sources.