Ireland's National Archives used OCR and robotic process automation to transcribe nearly 3 million 1926 Census entries — a task estimated at 23 person-years by hand — publishing the fully searchable records online in April 2026.
~3 million
Census entries transcribed
~45 million
Discrete data points
23 person-years
Estimated manual transcription effort
Details
Promoter
National Archives of Ireland
Period
2024-2026
Region (NUTS)
IE061
Keywords
public records, archives, digitisation, OCR, robotic process automation
Context
Ireland's 1926 Census - the Irish Free State's first - held almost 3 million individual entries (about 45 million discrete data points) in paper returns; manually transcribing them was estimated to take one person 23 years.
Objectives
To make the complete 1926 Census fully searchable online by combining OCR with robotic process automation, while preserving accuracy for difficult material such as entries in the seancló Gaelic typeface.
Activities
The National Archives digitised the original paper returns and applied an automated OCR/RPA transcription pipeline built on patterns first developed for the 1911 Census. Every transcribed field, including seancló entries, was checked twice by National Archives staff against the original scanned image before publication.
Results
The complete, searchable 1926 Census returns went live on the National Archives website on 18 April 2026, making a task estimated at 23 person-years of manual work feasible within a few years.
Conclusions
Combining automation with mandatory human double-verification allowed a large historical dataset to be published reliably and relatively quickly, though the practice is a one-off national archive project rather than a demonstrated transfer to other archives.
Implementation
Indicative cost
Medium (€50k–€500k)
Time to results
Medium (1–3 years) — Digitisation and transcription pipeline built 2024-2026; full searchable dataset published online 18 April 2026.
Staffing & skills
National Archives of Ireland staff manually double-checked every transcribed field against the original scanned image before publication, including entries in the seancló Gaelic typeface
Conditions for success
Reused and extended OCR/transcription patterns first developed for the 1911 Census rather than starting from scratch
Combined automation (OCR + RPA) with mandatory human double-verification of every field before publication
Where it fits
Governance type
national archival institution
Scale
national - entire 1926 Census (~3 million entries)
Income level
high-income
Commonly funded by
National / regional programmes
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Do you run this practice?
Claim it —
verified implementers get a public contact pathway and can propose corrections.
Data sources
Where this practice's information was retrieved from, and when.
Bank of Namibia's AI-assisted robotic-process-automation bot 'Onguvi' cut government payment-transaction processing to under two minutes, saving a reported N$6–7 million …