evidoria

← Back to browse

Good practice Imported

Texas STAAR Automated Essay Scoring — TEA Routes Every Written Answer Through an NLP Engine, With 25% Human Rescoring

United States of America · Austin · See the United States of America profile · See the Austin profile

Top 80% 33/100 · Ask Evidence Copilot about this practice

From spring 2024 the Texas Education Agency scores all open-ended STAAR answers with an NLP engine and rescores about 25% by hand, saving $15–20M a year. A spike in zero scores and equity concerns make it a cautionary case.

Texas STAAR Automated Essay Scoring — TEA Routes Every Written Answer Through an NLP Engine, With 25% Human Rescoring

Details

Promoter
Texas Education Agency (TEA)
Period
2023–2025
Keywords
K-12 assessment, automated scoring, natural language processing, state testing

Description

The Texas Education Agency began limited use of an automated scoring engine in December 2023 and applied it to constructed-response (open-ended) questions on the STAAR test across reading, writing, science and social studies from spring 2024. The engine uses natural language processing; computers score every response first, then roughly 25% of responses are rescored by humans. Responses the engine is not confident about, or that contain unfamiliar content such as slang or non-English text, are routed automatically to human scorers, and testing administrators review results daily and send random samples for human checking.
TEA reported annual savings of $15–20 million; the number of temporary scorers fell from about 6,000 in 2023 to fewer than 2,000 in 2024. Families can request rescoring for $50, waived if the score rises.
The rollout is contested. Reporting on the fall 2023 English II end-of-course exam found nearly 80% of written responses scored zero, versus roughly 25% in spring 2023 when humans graded. TEA attributed the difference to seasonal variation in test-takers (more retesters in the fall), while superintendents and teachers questioned whether the engine rewards originality and handles creative writing fairly, and coverage of the redesigned test noted that Black students had the highest share of zero scores. Evidence is descriptive and largely agency-reported; no independent audit of the engine's accuracy or bias was found in the cited sources, which is why the case is scored low on equity and transparency despite its clear scale and cost effect.

Read the full analysis: https://www.texastribune.org/2024/04/09/staar-artificial-intelligence-computer-grading-texas

Implementation

Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful