evidoria

← Back to browse

Good practice

NAPLAN Automated Essay Scoring Trial — ACARA's Evidence-Based Decision to Keep Human Markers

Australia · Sydney · See the Australia profile

ACARA found four AI essay-scoring systems scored NAPLAN writing as reliably as human markers, yet Australia's Education Council decided in Dec 2017 against AI marking for high-stakes national testing, citing bias and opacity risks.

1014 essays
Essays used to calibrate AI scoring systems (2015)
339 essays
Essays used to test AI scoring systems (2015)
4.5 million tests
Online NAPLAN tests completed (2026)
1.3 million students
Students who sat NAPLAN (2026)
9300 schools
Schools participating in NAPLAN (more than) (2026)
NAPLAN Automated Essay Scoring Trial — ACARA's Evidence-Based Decision to Keep Human Markers

Details

Maturity
Discontinued
Promoter
Australian Curriculum, Assessment and Reporting Authority (ACARA)
Period
2015 evaluation; Dec 2017 policy decision; ongoing as of 2026
Keywords
K-12 assessment, standardised testing, automated essay scoring, national exams, education policy

Context

In 2015, the Australian Curriculum, Assessment and Reporting Authority (ACARA) evaluated automated essay-scoring (AES) systems from four vendors — Measurement Incorporated, Pearson, Pacific Metrics and MetaMetrics — against NAPLAN persuasive-writing scripts marked by human assessors. The systems were calibrated on 1,014 essays and then tested on a further 339 essays. ACARA had targeted 2020 for a fully automated NAPLAN writing assessment, with human markers retained only as a backup check, positioning this evaluation as a step toward that goal.

Objectives

ACARA's evaluation aimed to determine whether AI scoring systems could match human markers closely enough, on both overall scores and individual writing criteria, to justify moving NAPLAN's high-stakes national writing assessment toward automation. Advocates such as education researcher Professor John Hattie argued automated marking could be far more accurate and considerably cheaper than human marking.

Activities

All four vendor systems were scored for agreement with human markers across overall scores and individual writing criteria using the calibration and test essay sets. Following the technical evaluation, teachers' unions raised sustained opposition, arguing that proprietary, non-public scoring algorithms could not reliably judge creativity, irony or argument logic, and risked bias against particular groups of students. In December 2017, Australia's Education Council, the national ministerial body overseeing NAPLAN, formally decided against using AES for NAPLAN writing, halting the path toward full automation.

Results

All four AI essay-scoring systems achieved score agreement with human markers that was comparable across both overall scores and individual writing criteria. Despite this, ACARA's plan to reach full automation by 2020 was abandoned after Australia's Education Council's December 2017 decision against AES for NAPLAN writing. Nearly a decade later, ACARA's official 27 March 2026 results release — covering roughly 4.5 million online tests taken by about 1.3 million students across more than 9,300 schools — notes that writing responses still take substantially longer to process and report than the other three tested domains, consistent with continued reliance on trained human markers.

Conclusions

This is a rare, well-documented case of an education system testing AI against a rigorous technical benchmark, finding it statistically competitive with human markers, and still choosing not to deploy it at the highest-stakes level. Australia's Education Council prioritised transparency, contestability and human judgement over the efficiency and cost gains that automation promised.

Implementation

Indicative cost
Medium (€50k–€500k)
Time to results
Long (> 3 years)

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful