The UK Foreign Office uses Microsoft AI Builder and GPT-4.1 to auto-classify consular correspondence, publishing UAT accuracy of 73-95% by task. Officials report an 80% drop in written enquiries within three months of its March 2024 launch; every reply still passes human review.
Reported drop in written enquiries after launch (within 3 months of March 2024 launch)
up to 50 %
Reported drop in phone calls after launch (within 3 months of March 2024 launch)
~30,500 emails/month
Average monthly mailbox volume since launch
60 days
Data retention period
Details
Maturity
Established
Promoter
Foreign, Commonwealth & Development Office (FCDO)
Period
2024–present
Keywords
Foreign affairs, consular services, correspondence management, generative AI
Context
The FCDO receives roughly 100,000 written consular enquiries a year from British nationals abroad — lost passports, deaths overseas, marriages, emergencies — and manual triage meant response times of several days. Correspondence Triage combines Microsoft AI Builder models (entity extraction, category classification, language detection), GPT-4.1 and GPT-4.1 mini prompts (summarisation, theme detection, sentiment) and robotic process automation that logs the result into the FCDO's case-management system, eCase.
Objectives
The tool assigns a case type and predicted fields for each incoming email; caseworkers can overwrite any prediction, and the system explicitly does not draft or send responses — every reply still goes through human review. Data is held for 60 days then deleted.
Results
The FCDO's published Algorithmic Transparency Record reports task-level accuracy from user-acceptance testing: 73.33% for auto email classification, an F1 score of 0.8512 for extracting geographical names, and 0.90-0.95 accuracy for case-type and 'no response required' determinations. Separately, FCDO officials told public-sector press that written enquiries fell by around 80% and phone calls by up to 50% within three months of the March 2024 go-live, with the mailbox since averaging roughly 30,500 emails a month.
Conclusions
Those outcome figures are self-reported by the department rather than independently audited, and the ATR itself flags real risks: dependency on the system if it goes offline, and Microsoft's content-moderation layer filtering some sensitive inputs, with flagged cases going to manual review. The team deliberately avoided a free-form generative interface in favour of matching enquiries to pre-approved, human-written content.
Implementation
Indicative cost
Medium (€50k–€500k) — Built on existing Microsoft AI Builder and GPT-4.1/GPT-4.1 mini models plus RPA integrated into the existing eCase case-management system; no separate cost figure published.
Time to results
Medium (1–3 years) — Live since March 2024; UAT and outcome figures reported through January 2025 press coverage.
Staffing & skills
Foreign, Commonwealth & Development Office (FCDO) caseworkers, FCDO digital/AI team, Microsoft (AI Builder, GPT-4.1 models)
Conditions for success
Human review of every AI-generated reply before sending
Caseworkers able to overwrite any AI prediction
Defined data retention limit (60 days)
Avoiding a free-form generative interface in favour of matching to pre-approved human-written content
Common failure modes
Outcome figures (enquiry/call volume drops) are self-reported by the department, not independently audited
The UK's Government Digital Service piloted GOV.UK Chat, a retrieval-augmented generative-AI assistant answering citizens' tax, benefits and visa questions — …