evidoria

← Back to browse

Good practice Imported

Swedish Lion — Riksarkivet's AI Model Making a Million Handwritten Historical Documents Searchable

Sweden · Stockholm · See the Sweden profile · See the Stockholm profile

Evidence: Descriptive / self-reported Top 14% 80/100 · Ask Evidence Copilot about this practice

Sweden's National Archives built an in-house AI handwriting-recognition model, trained via a three-year citizen-science transcription effort, to make over a million pages of 17th- and 18th-century court and witchcraft-trial records full-text searchable, and released the results a

~1.2 million pages
Pages made full-text searchable (as of the practice's reported period)
1 million+ documents
Initial batch of handwritten documents processed (initial batch)
Swedish Lion — Riksarkivet's AI Model Making a Million Handwritten Historical Documents Searchable

Details

Maturity
Scaling
Promoter
Riksarkivet (Swedish National Archives)
Period
2022–2025
Region (NUTS)
SE11
Keywords
cultural heritage, archives, digitization, handwritten text recognition

Context

Sweden's National Archives (Riksarkivet) operates its own AI research unit, AIRA, which built 'Swedish Lion', a handwritten text recognition (HTR) model trained specifically to read Swedish handwriting from roughly 1600 to 1900. The training data came from 'Transkriberingsnod Sverige', a three-year, publicly funded collaboration in which volunteer genealogy researchers manually transcribed historical documents to build the ground-truth dataset the model needed.

Objectives

The project aims to make large historical archival collections full-text searchable, starting with two major collections: the 17th-century Witchcraft Commission archive and the 18th-century Svea Court of Appeal protocols.

Activities

Riksarkivet has published the machine-interpreted texts as open data, downloadable via API or as XML files, and previously open-sourced its underlying HTR/OCR tooling, including the 'HTRflow' framework, for reuse by other archives and researchers.

Results

The model has made roughly 1.2 million pages full-text searchable in the National Archives Database, from an initial batch of over one million handwritten documents. Riksarkivet has said millions more archival pages will be processed and published progressively through 2025.

Conclusions

Because the project's own materials do not publish a formal, independently verified accuracy or error-rate figure, the practice is evidenced primarily by its delivered, publicly checkable output rather than by a benchmarked accuracy score.

Implementation

Indicative cost
Medium (€50k–€500k)
Time to results
Long (> 3 years) — Training data was built over a three-year citizen-science transcription project; the model has since been applied to two major historical collections, with further pages planned for release through 2025.
Staffing & skills
Riksarkivet's in-house AI research unit (AIRA), Volunteer genealogy researchers participating in Transkriberingsnod Sverige (Transcription Node Sweden)

Conditions for success

  • A large, high-quality ground-truth transcription dataset built through a multi-year citizen-science effort
  • Publishing outputs and tooling (HTRflow) as open data and open source to enable outside verification and reuse

Common failure modes

  • No independently verified accuracy or error-rate figure has been published for the HTR model itself.

Where it fits

Governance type
national archive / cultural heritage institution
Scale
national
Income level
high-income

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful