Facing a mandatory 25-year declassification deadline and a year of manual review per annual cohort, the U.S. State Department trained an ML model that confidently decided over 60% of 121,536 diplomatic cables in minutes, matching human reviewers 97-99% of the time.
121536 cables
Diplomatic cables processed in the 1998 cohort (2023)
60 % of cohort (>72,000 cables)
Cables confidently decided by the model (2023)
97-99 %
Agreement with human reviewers on validation set (2023)
99.29 %
Accuracy on declassification calls (independent account) (2023)
81.43 %
Accuracy on withholding calls (independent account) (2023)
400000 USD
Development cost (2022-2023)
47218 cables (39% of cohort)
Cables still requiring manual human review (2023)
Details
Maturity
Scaling
Promoter
U.S. Department of State — Office of Information Programs and Services (Declassification Division) / Center for Analytics
Period
2022-2024
Keywords
records management, declassification, government transparency, machine learning, FOIA
Context
Executive Order 13526 requires U.S. government agencies to review classified records for automatic declassification once they turn 25 years old, and the volume of State Department diplomatic cables has grown roughly six-fold since the 1990s, making full manual review increasingly unsustainable.
Objectives
From October 2022, the Declassification Division, the Bureau of Information Resource Management and the Center for Analytics set out to train a supervised machine-learning model on past human declassification decisions to triage cables into 'confidently declassify', 'confidently exempt' and 'needs human review' bands, reducing the volume requiring full manual review.
Activities
The model was trained on human declassification decisions from 1995-1997 cables. After a successful pilot, it was redesignated a full Program in 2023 and applied to the 1998 cohort of 121,536 cables, producing a confidence score and band per cable.
Results
The model confidently decided over 72,000 cables (about 60% of the cohort) in 20-30 minutes of computing time each, versus roughly a year of manual review previously. The State Department reports 97-99% agreement with human reviewers on the validation set; an independent account by the American Historical Association gives more granular figures of 99.29% accuracy on declassification calls versus 81.43% on withholding calls, and a development cost of about $400,000. About 39% of cables (47,218 of 121,536) still required manual human review because the model could not confidently classify them.
Conclusions
The Department describes the model as deliberately 'overprotective', erring toward withholding when uncertain, and cautions against generalising results to other record types without retraining; separate pilots for AI-assisted FOIA search remain in early testing, not full deployment.
Implementation
Indicative cost
Low (< €50k) — Development cost reported at approximately $400,000 (American Historical Association account); no separate operating-cost figures published.
Time to results
Short (< 1 year) — Model training began October 2022 on a pilot basis; redesignated a full Program in 2023 and applied to the 121,536-cable 1998 cohort, each confidently-decided cable processed in 20-30 minutes of computing time versus roughly a year of manual review previously.
Staffing & skills
Declassification Division, Bureau of Information Resource Management, Center for Analytics
Conditions for success
Training on a large corpus of prior human declassification decisions (1995-1997 cables) before deployment
Deliberately calibrating the model to be 'overprotective' (erring toward withholding) when confidence is low
Retaining mandatory human review for the ~39% of cases the model cannot confidently classify
Where it fits
Governance type
national government agency
Scale
national
Income level
high_income
Commonly funded by
National / regional programmes
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Do you run this practice?
Claim it —
verified implementers get a public contact pathway and can propose corrections.
Data sources
Where this practice's information was retrieved from, and when.