A CHI 2026 randomised trial at the University of Michigan found AI-drafted feedback, reviewed by teaching assistants, improved 354 economics students' essay revisions by an effect comparable to moving from the 50th to 70th percentile; TAs kept final control.
354 students
Students in study
11 TAs
Teaching assistants involved
88 %
AI suggestions used by TAs without changes
12 %
AI suggestions edited or rejected by TAs
50th to 70th percentile
Effect size of AI-mediated feedback on revision quality
Researchers at the University of Michigan built FeedbackWriter, a tool that drafts feedback suggestions for teaching assistants marking student essays, tested in a large introductory economics course with 354 students and 11 teaching assistants. The study was funded by the US National Science Foundation and presented at the ACM CHI 2026 conference.
Objectives
The trial set out to test whether AI-drafted feedback, reviewed and editable by teaching assistants before reaching students, could improve the quality of student essay revisions compared with feedback from teaching assistants alone.
Activities
In the randomised controlled trial, TAs could accept, edit or discard the AI's draft feedback for each of two knowledge-intensive essay assignments before it reached students.
Results
Essays that received AI-mediated feedback led to higher-quality student revisions than essays receiving feedback from TAs alone, an effect the authors describe as roughly equivalent to moving a student from the 50th to the 70th percentile. TAs agreed with and used 88% of the AI's suggested judgments without changes, editing or rejecting the remaining 12%.
Conclusions
The study was conducted in a single course at one institution, so its generalisability to other subjects, class sizes and institutions is not yet established.
Implementation
Indicative cost
Low (< €50k)
Time to results
Short (< 1 year)
Staffing & skills
11 teaching assistants, Research team: Xinyi Lu, Kexin Phyllis Ju, Mitchell Dudley, Larissa Sano, Xu Wang (University of Michigan)
Conditions for success
Teaching assistants review, edit or discard every AI-drafted suggestion before it reaches students (human-in-the-loop)
NSF funding supporting a dedicated research and tool-development team
Common failure modes
Conducted in a single course at one institution; generalisability to other subjects, class sizes and institutions not yet established
No equity subgroup analysis reported
Commonly funded by
National / regional programmes
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Replication kit
Reusable artefacts from this practice — as published by their sources.