evidoria

← Back to browse

Good practice Imported

Cambridge Study Finds AI Not Yet Reliable for Grading University Essays

United Kingdom · Cambridge · See the United Kingdom profile · See the Cambridge profile

Top 72% 40/100 · Ask Evidence Copilot about this practice

Cambridge psychologists and AI researchers tested Claude and ChatGPT grading 750 real essays from three UK universities. The systems matched human examiners' grade bands only 35–65% of the time, rewarding essay length and vocabulary over academic substance.

Cambridge Study Finds AI Not Yet Reliable for Grading University Essays

Details

Promoter
University of Cambridge
Period
2026
Keywords
higher education, psychology, assessment research

Description

A University of Cambridge-led team of psychologists and AI researchers evaluated three frontier large language models — including recent versions of Claude and ChatGPT (as of April 2026) — on their ability to grade real undergraduate work. The team assembled a corpus of more than 750 psychology essays submitted at three UK universities and compared AI-assigned grade bands (first, 2:1, 2:2, and so on) against the grades originally awarded by human examiners.
The AI systems matched the human examiners' grade band only 35–65% of the time depending on the model and essay set — far from the reliability required for high-stakes assessment. The researchers identified a specific failure mode: all three systems were oversensitive to surface linguistic features, systematically over-rewarding essay length and vocabulary range regardless of whether the underlying argument was strong. AI graders also routinely undervalued essays that human markers had rated as top-scoring, and overvalued some of the weakest submissions.
When asked to generate written feedback rather than just a grade, the AI systems produced comments three to eight times longer than those written by the original human assessors — verbose but not necessarily more useful to students. The Cambridge team's conclusion was explicitly cautionary: AI is 'not yet good enough' to grade open-ended university essays without significant human oversight, a finding intended to slow down universities considering automated grading before more rigorous safeguards and validation are in place.

Read the full analysis: https://www.cam.ac.uk/stories/ai-university-essay-grading

Implementation

Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful