Estonia's AI-Assisted Grading Trial for National School-Leaving Essay Exams
Estonia
A 2026 peer-reviewed study tested LLM and statistical-NLP grading of Estonia's national school-leaving essay exams across two full cohorts. AI …
Australia · Sydney · See the Australia profile
ACARA found four AI essay-scoring systems scored NAPLAN writing as reliably as human markers, yet Australia's Education Council decided in Dec 2017 against AI marking for high-stakes national testing, citing bias and opacity risks.
In 2015, the Australian Curriculum, Assessment and Reporting Authority (ACARA) evaluated automated essay-scoring (AES) systems from four vendors — Measurement Incorporated, Pearson, Pacific Metrics and MetaMetrics — against NAPLAN persuasive-writing scripts marked by human assessors. The systems were calibrated on 1,014 essays and then tested on a further 339 essays. ACARA had targeted 2020 for a fully automated NAPLAN writing assessment, with human markers retained only as a backup check, positioning this evaluation as a step toward that goal.
ACARA's evaluation aimed to determine whether AI scoring systems could match human markers closely enough, on both overall scores and individual writing criteria, to justify moving NAPLAN's high-stakes national writing assessment toward automation. Advocates such as education researcher Professor John Hattie argued automated marking could be far more accurate and considerably cheaper than human marking.
All four vendor systems were scored for agreement with human markers across overall scores and individual writing criteria using the calibration and test essay sets. Following the technical evaluation, teachers' unions raised sustained opposition, arguing that proprietary, non-public scoring algorithms could not reliably judge creativity, irony or argument logic, and risked bias against particular groups of students. In December 2017, Australia's Education Council, the national ministerial body overseeing NAPLAN, formally decided against using AES for NAPLAN writing, halting the path toward full automation.
All four AI essay-scoring systems achieved score agreement with human markers that was comparable across both overall scores and individual writing criteria. Despite this, ACARA's plan to reach full automation by 2020 was abandoned after Australia's Education Council's December 2017 decision against AES for NAPLAN writing. Nearly a decade later, ACARA's official 27 March 2026 results release — covering roughly 4.5 million online tests taken by about 1.3 million students across more than 9,300 schools — notes that writing responses still take substantially longer to process and report than the other three tested domains, consistent with continued reliance on trained human markers.
This is a rare, well-documented case of an education system testing AI against a rigorous technical benchmark, finding it statistically competitive with human markers, and still choosing not to deploy it at the highest-stakes level. Australia's Education Council prioritised transparency, contestability and human judgement over the efficiency and cost gains that automation promised.
Where this practice's information was retrieved from, and when.
Estonia
A 2026 peer-reviewed study tested LLM and statistical-NLP grading of Estonia's national school-leaving essay exams across two full cohorts. AI …
South Korea
Seoul-based Riiid Inc. applies deep reinforcement learning to adaptive TOEIC prep: score predicted in 6–10 questions; 2.5 million users (TechCrunch …
Portugal
Portugal's 2026 rollout of digital correction for national secondary exams (81,000+ students) hit widespread login failures, lost and duplicated scripts, …
Australia
Instead of an unwinnable AI-detection arms race, Sydney redesigned assessment into two lanes — secure in-person assessment of core capability, …
Open full copilot Grounded in cited practices — always check the sources.