evidoria

← Back to browse

Good practice Imported

PNLD Image-Quality Triage — A Randomised Trial of a CNN Assistant With 76 of Brazil's Textbook Analysts

Brazil · Brasília · See the Brazil profile · See the Brasília profile

Evidence: Randomised controlled trial Top 25% 73/100 · Ask Evidence Copilot about this practice

Brazil's PNLD textbook programme tested a CNN sorting book images as sharp, defocus-blurred or motion-blurred in a two-arm 1:1 randomised study with 76 serving analysts. The AI arm assessed substantially more images with quality preserved; pre-to-post gains were modest.

76 analysts
Textbook analysts randomised (1:1, AI-assisted vs control) (2025-2026)
PNLD Image-Quality Triage — A Randomised Trial of a CNN Assistant With 76 of Brazil's Textbook Analysts

Details

Maturity
Pilot
Promoter
Fundo Nacional de Desenvolvimento da Educação (FNDE) with NEES/Universidade Federal de Alagoas
Period
2025–2026
Keywords
public administration, education procurement, document review, computer vision, evaluation, randomised trial

Context

The Programa Nacional do Livro e do Material Didático (PNLD), created in 1937 and run by the Fundo Nacional de Desenvolvimento da Educação (FNDE) with the Ministry of Education, evaluates and distributes textbooks to Brazil's public schools. Its assessment cycle can take at least two years and involves hundreds of professionals doing manual review work, including judging the technical quality of images printed in candidate textbooks.

Objectives

Researchers at the Nucleus of Excellence in Social Technologies (NEES) at the Federal University of Alagoas, with co-authors at several Brazilian and international universities, built a convolutional neural network to sort each textbook image as sharp, defocused-blurred or motion-blurred, so that analysts can concentrate on the cases the model flags.

Activities

The team ran a parallel two-arm study with 1:1 randomisation and 76 textbook analysts, people who actually do PNLD assessment work rather than a convenience sample. Participants took a baseline test and a post-test after the initial assessment; the primary outcome was the post-test score, and secondary outcomes were the analysts' productivity and the quality of their assessments.

Results

The AI component significantly increased productivity: the experimental group assessed substantially more images than the control group, and assessment quality was preserved rather than traded away for speed. The primary outcome was weaker, the difference between pre-test and post-test was modest, though the authors read the direction as evidence that working with the AI component improved analysts' own ability to tell image-defect categories apart.

Conclusions

The trial tests one narrow component, blur classification on images, not the pedagogical evaluation at the heart of PNLD, and the productivity finding should not be read as a claim about the assessment cycle as a whole. A randomised study of a component is also not the same as evidence of a production rollout; FNDE and MEC have a broader PNLD modernisation effort under way with NEES, but this specific trial establishes efficacy, not deployment at scale. Randomised evaluation of a government AI tool, with the actual public officials who use it, published open access including a null-ish primary outcome, is rare among public-sector AI productivity claims.

Implementation

Indicative cost
Low (< €50k) — Supported by FNDE under TED 10320 and by Brazil's CNPq through the INCT IA.Edu institute (grant 408483/2024-5).
Time to results
Short (< 1 year) — Trial conducted in 2025-2026 and published in Scientific Reports on 10 August 2026; PNLD's own multi-year modernisation effort with NEES continues alongside it.
Staffing & skills
Fundo Nacional de Desenvolvimento da Educação (FNDE) with the Ministry of Education, NEES (Nucleus of Excellence in Social Technologies) at the Federal University of Alagoas, Co-authors at the Federal University of the Agreste of Pernambuco, the Federal Rural University of Pernambuco, the Harvard Graduate School of Education and the University of Pennsylvania, 76 serving PNLD textbook analysts as trial participants

Conditions for success

  • A randomised, controlled design run with actual public officials who do the assessment work, not a convenience sample
  • Baseline and post-test measurement of both productivity and assessment quality, not speed alone
  • Publishing the trial open access, including a modest primary-outcome result, rather than only favourable findings

Common failure modes

  • The trial tests one narrow component, image blur classification, not the pedagogical evaluation at the heart of PNLD.
  • A randomised trial of a component is not the same as evidence of a full production rollout.
  • No algorithmic governance documentation, appeal route, or production oversight arrangement has been published.

Where it fits

Governance type
national education programme with a university research partnership
Scale
component-level trial (76 analysts)
Income level
upper-middle-income

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Replication kit

Reusable artefacts from this practice — as published by their sources.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful