evidoria

← Back to browse

Good practice

Estonia's AI-Assisted Grading Trial for National School-Leaving Essay Exams

Estonia · Tallinn · See the Estonia profile

A 2026 peer-reviewed study tested LLM and statistical-NLP grading of Estonia's national school-leaving essay exams across two full cohorts. AI scores fell within the human-rater range in 60% of cases, with final decisions kept with human assessors.

60 %
AI grades falling within human-rater agreement range (trial cohorts)
Estonia's AI-Assisted Grading Trial for National School-Leaving Essay Exams

Details

Maturity
Pilot
Promoter
Estonian Education and Youth Board (Harno), with Tallinn University researchers
Period
2024–2026 pilot; national e-exam rollout planned from 2027
Keywords
national exams, NLP, LLM grading, assessment pilot

Context

Estonia's Education and Youth Board (Harno) is digitising the country's basic-school and upper-secondary exit exams, with fully digital e-exams planned from 2027. Ahead of that shift, Harno worked with researchers including a team from Tallinn University to test whether large language models and statistical NLP could support grading of the national mother-tongue and school-leaving essay exams.

Results

The study compared machine-generated scores against official human panel scores on two full national cohorts of trial essays, using the same curriculum-based rubric used by human graders. In 60% of cases the language model's grade fell within the range that human assessors vary among themselves, with the strongest agreement on more objective criteria such as correct use of source texts. The study also tested the models for bias and prompt-injection vulnerabilities.

Conclusions

Because EU rules bar machines from making binding decisions on student exams, any deployment keeps a human assessor as final decision-maker, with AI treated as decision support (e.g. flagging source-text usage) rather than a replacement grader. Harno was still reviewing final results at time of reporting, and subjective criteria such as argument quality remain a human judgement call.

Implementation

Indicative cost
Medium (€50k–€500k) — No public budget figure found; cost_band set to medium, reflecting a national government research collaboration building toward system-wide digital exams.
Time to results
Medium (1–3 years) — Pilot ran 2024–2026 with full e-exam rollout planned from 2027; timeline_band set to medium.
Staffing & skills
Human assessors remain the final decision-makers, as required by EU rules, Tallinn University researchers built and tested the models

Conditions for success

  • Use of the same curriculum-based rubric as human graders
  • Explicit testing for bias and prompt-injection vulnerabilities
  • AI positioned as decision support, not a replacement for human judgement

Common failure modes

  • Subjective criteria such as argument quality still require human judgement
  • Final results still under review at time of reporting

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful