evidoria

← Back to browse

Good practice Imported

Anytime Testing Machine (ATM) — Pratham and Anthropic's Claude-Powered Formative Assessment in Indian Schools

India · Mumbai · See the India profile

Pratham and Anthropic piloted the Claude-powered Anytime Testing Machine, which grades handwritten answers and gives personalised feedback, lifting accuracy vs. human graders from ~30% to ~80% across 1,500 students and 5,000+ women in Pratham's Second Chance programme.

~30 %
Grading-accuracy alignment with human graders (before refinement) (2026 pilot)
~80 %
Grading-accuracy alignment with human graders (after refinement) (2026 pilot)
~90 %
Question-generation alignment with Bloom's Taxonomy benchmarks (2026 pilot)
1500 students
Students in initial pilot (2026)
Anytime Testing Machine (ATM) — Pratham and Anthropic's Claude-Powered Formative Assessment in Indian Schools

Details

Maturity
Pilot
Promoter
Pratham Education Foundation, with Anthropic
Period
2026-present
Keywords
education technology, assessment, nonprofit

Context

In February 2026, Anthropic announced its first strategic AI-lab partnership with an Indian education nonprofit, Pratham Education Foundation, to co-develop the Anytime Testing Machine (ATM), a Claude-powered, end-to-end formative-assessment system for Indian schools.

Objectives

The system aims to grade handwritten student answers against curriculum-aligned rubrics and generate personalised feedback in Hindi and English, with teachers reviewing and approving all AI-generated feedback before it reaches students.

Activities

Students handwrite answers, photograph them, and the system converts the images to text and evaluates responses using an LLM-as-a-judge approach. The tool was piloted with 1,500 students across 20 schools and adapted for Pratham's Second Chance programme, which supports over 5,000 women preparing for India's grade 10 board exam.

Results

Iterative refinement raised the system's grading-accuracy alignment with expert human graders from roughly 30% to roughly 80%, and question-generation accuracy reached about 90% alignment with Bloom's Taxonomy benchmarks. These are grading- and question-quality metrics, not verified gains in student learning, and no independent learning-outcome evaluation has yet been published.

Conclusions

Anthropic and Pratham plan to expand ATM to 100 schools and the Second Chance track to 15,000 women by the end of 2026, and are designing a randomised controlled trial under the Teaching at the Right Level (TaRL) methodology to measure learning impact.

Implementation

Indicative cost
Medium (€50k–€500k) — Funded through Anthropic's strategic partnership with Pratham; no public budget figures, but involves proprietary LLM API usage and iterative development cycles.
Time to results
Short (< 1 year) — Piloted from February 2026; expansion to 100 schools and 15,000 Second Chance learners targeted by end of 2026, with an RCT under the TaRL methodology in the design stage.
Staffing & skills
Pratham Education Foundation, Anthropic (technical/AI partner), Teachers (review and approve all AI-generated feedback before students see it)

Conditions for success

  • Human-in-the-loop teacher review of all AI-generated feedback before release to students
  • Iterative refinement cycles that measurably improved grading-accuracy alignment with human graders
  • Existing Pratham Second Chance programme infrastructure to reach women preparing for board exams

Common failure modes

  • Initial grading-accuracy alignment with human graders was only about 30%, showing early LLM-graded assessment can be unreliable without iterative human correction
  • No independent learning-outcome evaluation exists yet; a planned RCT has not been completed

Where it fits

Governance type
nonprofit + private AI-lab partnership
Scale
pilot (20 schools, 1,500 students; Second Chance track scaling toward 15,000 women)
Income level
lower-middle-income

Commonly funded by

Philanthropic / foundation funding

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful

★ 33

Turnitin AI writing detection

United States of America

Widely deployed AI-writing detection — but documented false positives, bias against non-native English writers, and wrongful cheating accusations make it …