Letrus AI Writing Feedback — Closing Brazil's ENEM Essay Gap in Espírito Santo
Brazil
Espírito Santo's public schools used the Letrus AI platform to give ENEM essay writers rapid feedback; a J-PAL randomized evaluation …
United States of America · Austin · See the United States of America profile · See the Austin profile
Top 80% 33/100 · Ask Evidence Copilot about this practice
From spring 2024 the Texas Education Agency scores all open-ended STAAR answers with an NLP engine and rescores about 25% by hand, saving $15–20M a year. A spike in zero scores and equity concerns make it a cautionary case.
The Texas Education Agency began limited use of an automated scoring engine in December 2023 and applied it to constructed-response (open-ended) questions on the STAAR test across reading, writing, science and social studies from spring 2024. The engine uses natural language processing; computers score every response first, then roughly 25% of responses are rescored by humans. Responses the engine is not confident about, or that contain unfamiliar content such as slang or non-English text, are routed automatically to human scorers, and testing administrators review results daily and send random samples for human checking.
TEA reported annual savings of $15–20 million; the number of temporary scorers fell from about 6,000 in 2023 to fewer than 2,000 in 2024. Families can request rescoring for $50, waived if the score rises.
The rollout is contested. Reporting on the fall 2023 English II end-of-course exam found nearly 80% of written responses scored zero, versus roughly 25% in spring 2023 when humans graded. TEA attributed the difference to seasonal variation in test-takers (more retesters in the fall), while superintendents and teachers questioned whether the engine rewards originality and handles creative writing fairly, and coverage of the redesigned test noted that Black students had the highest share of zero scores. Evidence is descriptive and largely agency-reported; no independent audit of the engine's accuracy or bias was found in the cited sources, which is why the case is scored low on equity and transparency despite its clear scale and cost effect.
Read the full analysis: https://www.texastribune.org/2024/04/09/staar-artificial-intelligence-computer-grading-texas
Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.
Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.
Where this practice's information was retrieved from, and when.
Brazil
Espírito Santo's public schools used the Letrus AI platform to give ENEM essay writers rapid feedback; a J-PAL randomized evaluation …
Australia
ACARA found four AI essay-scoring systems scored NAPLAN writing as reliably as human markers, yet Australia's Education Council decided in …
China
Pigai is China's dominant AI essay-feedback platform, used by thousands of schools and tens of millions of EFL students since …
China
iFLYTEK's speech-recognition AI scores spoken English in China's zhongkao and gaokao oral exams across dozens of provinces, processing millions of …
Open full copilot Grounded in cited practices — always check the sources.