evidoria

← Back to browse

Good practice Imported

Ofqual's 2020 Exam-Grading Algorithm — A National Cautionary Case in Automated Assessment

United Kingdom · Coventry · See the United Kingdom profile

When COVID-19 cancelled exams, England's regulator Ofqual used a statistical model to standardise A-level and GCSE grades; it downgraded 39% of teacher-predicted grades, hit disadvantaged pupils hardest, and was scrapped within four days.

Ofqual's 2020 Exam-Grading Algorithm — A National Cautionary Case in Automated Assessment

Details

Promoter
Ofqual (Office of Qualifications and Examinations Regulation)
Period
2020
Keywords
algorithmic assessment, exam regulation, statistical moderation, education policy

Description

When the COVID-19 pandemic cancelled A-level and GCSE exams in England in 2020, the exams regulator Ofqual needed a way to award grades without sitting exams. Schools submitted Centre Assessment Grades (teachers' predicted grades) and a rank order of students within each grade. Ofqual then applied a statistical model that adjusted these predictions using each school's historical grade distribution from the previous three years and subject-level GCSE attainment data, in order to prevent nationwide grade inflation relative to prior years.

When results were released on 13 August 2020, the model had downgraded 39% of teachers' predicted grades by at least one grade. Ofqual's own deputy director confirmed that pupils from disadvantaged backgrounds were more likely to see larger downward adjustments, while the model was not applied to cohorts smaller than 15 students — a category that included many fee-paying schools — so the proportion of private-school pupils awarded top grades rose more than for state-school pupils. Public anger, student protests and cross-party criticism followed; on 17 August, just four days after results day, Ofqual reversed course and awarded students their teacher-predicted grades instead. Ofqual's chief regulator resigned on 25 August, and the Department for Education's permanent secretary resigned on 28 August.

The episode is now widely used in education-policy and data-ethics teaching as a case study in what can go wrong when an algorithm optimises for a population-level statistical target (matching historical national grade distributions) at the expense of individual fairness, without adequate pre-release scrutiny, affected-student consultation, or transparency about how the model would treat small and disadvantaged cohorts.

Read the full analysis: https://www.bera.ac.uk/blog/the-great-algorithm-fiasco

Implementation

Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful