New Zealand's NZQA trained an AI model on tens of thousands of past-marked scripts to grade the national NCEA Writing literacy exam, reaching 80% agreement with human markers and cutting results turnaround by 3.5 weeks for 55,000+ students in 2025.
Details
Promoter
New Zealand Qualifications Authority (NZQA)
Period
2024–present
Keywords
government agency, national qualifications authority, secondary education, assessment
Description
New Zealand Qualifications Authority (NZQA) built an Automated Text Scoring tool, training a machine-learning model on thousands of previously human-marked scripts to grade the NCEA co-requisite Writing standard — a national literacy assessment that most secondary students must pass to gain NCEA qualifications. The AI does the primary marking; human markers are retained only to check the most borderline scripts. In a 2024 trial covering 36,000 writing samples, NZQA reported the tool was "as reliable as human markers." Chief executive Grant Klinkum stated the model reached an 80% agreement rate with human markers. The tool was rolled out nationally for the co-requisite Writing exam sat in May 2025 by more than 55,000 students. About 20,000 scripts (roughly 40%, concentrated near the pass/fail boundary) were still checked by a human marker, and wherever AI and human scores disagreed, the human mark was used. Schools received results 3.5 weeks earlier than in the previous year. The evidence base covers grading accuracy and turnaround time, not direct student learning outcomes. The 20% disagreement rate observed in the trial is non-trivial for a high-stakes qualification gate, which is why NZQA retained mandatory human review for borderline cases rather than fully automating the exam.
Read the full analysis: https://www.1news.co.nz/2025/04/10/ai-to-mark-year-10-students-writing-tests/
Implementation
Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.
Do you run this practice?
Claim it —
verified implementers get a public contact pathway and can propose corrections.
Data sources
Where this practice's information was retrieved from, and when.
Universitas Kuningan researchers built and validated AKM Online, an AI-enhanced version of Indonesia's national competency assessment, testing 552 students across …
An RCT across four political science courses at Houston and partner universities found GPT-4-assisted grading of short-answer questions, using instructor …