NAPLAN Automated Essay Scoring Trial — ACARA's Evidence-Based Decision to Keep Human Markers
Australia
ACARA found four AI essay-scoring systems scored NAPLAN writing as reliably as human markers, yet Australia's Education Council decided in …
Estonia · Tallinn · See the Estonia profile · See the Tallinn profile
Evidence: Observational / pre–post Top 42% 53/100 · Ask Evidence Copilot about this practice
A 2026 peer-reviewed study tested LLM and statistical-NLP grading of Estonia's national school-leaving essay exams across two full cohorts. AI scores fell within the human-rater range in 60% of cases, with final decisions kept with human assessors.
Estonia's Education and Youth Board (Harno) is digitising the country's basic-school and upper-secondary exit exams, with fully digital e-exams planned from 2027. Ahead of that shift, Harno worked with researchers including a team from Tallinn University to test whether large language models and statistical NLP could support grading of the national mother-tongue and school-leaving essay exams.
The study compared machine-generated scores against official human panel scores on two full national cohorts of trial essays, using the same curriculum-based rubric used by human graders. In 60% of cases the language model's grade fell within the range that human assessors vary among themselves, with the strongest agreement on more objective criteria such as correct use of source texts. The study also tested the models for bias and prompt-injection vulnerabilities.
Because EU rules bar machines from making binding decisions on student exams, any deployment keeps a human assessor as final decision-maker, with AI treated as decision support (e.g. flagging source-text usage) rather than a replacement grader. Harno was still reviewing final results at time of reporting, and subjective criteria such as argument quality remain a human judgement call.
National / regional programmes
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.
Where this practice's information was retrieved from, and when.
Australia
ACARA found four AI essay-scoring systems scored NAPLAN writing as reliably as human markers, yet Australia's Education Council decided in …
Portugal
Portugal's 2026 rollout of digital correction for national secondary exams (81,000+ students) hit widespread login failures, lost and duplicated scripts, …
Australia
Instead of an unwinnable AI-detection arms race, Sydney redesigned assessment into two lanes — secure in-person assessment of core capability, …
Croatia
CARNET's ESF+-funded BrAIn project piloted Croatia's first AI curriculum in VET schools, covering responsible AI use, plagiarism and authorship crediting; …
Open full copilot Grounded in cited practices — always check the sources.