evidoria

← Back to browse

Good practice Imported

FeedbackWriter — Randomised Trial of AI-Mediated Feedback for Teaching Assistants at the University of Michigan

United States of America · Ann Arbor · See the United States of America profile

Evidence: Randomised controlled trial Top 43% 53/100 · Ask Evidence Copilot about this practice

A CHI 2026 randomised trial at the University of Michigan found AI-drafted feedback, reviewed by teaching assistants, improved 354 economics students' essay revisions by an effect comparable to moving from the 50th to 70th percentile; TAs kept final control.

354 students
Students in study
11 TAs
Teaching assistants involved
88 %
AI suggestions used by TAs without changes
12 %
AI suggestions edited or rejected by TAs
50th to 70th percentile
Effect size of AI-mediated feedback on revision quality

Details

Promoter
University of Michigan
Period
2025–2026 (presented at ACM CHI 2026)
Keywords
Higher education, teaching assistants, writing feedback, economics

Context

Researchers at the University of Michigan built FeedbackWriter, a tool that drafts feedback suggestions for teaching assistants marking student essays, tested in a large introductory economics course with 354 students and 11 teaching assistants. The study was funded by the US National Science Foundation and presented at the ACM CHI 2026 conference.

Objectives

The trial set out to test whether AI-drafted feedback, reviewed and editable by teaching assistants before reaching students, could improve the quality of student essay revisions compared with feedback from teaching assistants alone.

Activities

In the randomised controlled trial, TAs could accept, edit or discard the AI's draft feedback for each of two knowledge-intensive essay assignments before it reached students.

Results

Essays that received AI-mediated feedback led to higher-quality student revisions than essays receiving feedback from TAs alone, an effect the authors describe as roughly equivalent to moving a student from the 50th to the 70th percentile. TAs agreed with and used 88% of the AI's suggested judgments without changes, editing or rejecting the remaining 12%.

Conclusions

The study was conducted in a single course at one institution, so its generalisability to other subjects, class sizes and institutions is not yet established.

Implementation

Indicative cost
Low (< €50k)
Time to results
Short (< 1 year)
Staffing & skills
11 teaching assistants, Research team: Xinyi Lu, Kexin Phyllis Ju, Mitchell Dudley, Larissa Sano, Xu Wang (University of Michigan)

Conditions for success

  • Teaching assistants review, edit or discard every AI-drafted suggestion before it reaches students (human-in-the-loop)
  • NSF funding supporting a dedicated research and tool-development team

Common failure modes

  • Conducted in a single course at one institution; generalisability to other subjects, class sizes and institutions not yet established
  • No equity subgroup analysis reported

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Replication kit

Reusable artefacts from this practice — as published by their sources.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful