evidoria

← Back to browse

Good practice Imported

Woogle — the Dutch NLP Search Engine for 3 Million Freedom-of-Information Files, Now Being Tested on Estonia's Records

Netherlands · Amsterdam · See the Netherlands profile · See the Amsterdam profile

Evidence: Descriptive / self-reported Top 14% 80/100 · Ask Evidence Copilot about this practice

Woogle, built at the University of Amsterdam, uses OCR and document-classification models to index over 3 million Dutch FOIA dossiers (11M+ pages) from 1,000+ public bodies — and peer-reviewed research shows the same pipeline can be adapted to Estonia's public records.

3,000,000+
FOIA dossiers indexed
11,000,000+
Standardised pages indexed
1,000+
Dutch public bodies covered
~10,000
Estonian documents successfully uploaded in transfer test (2025)
Woogle — the Dutch NLP Search Engine for 3 Million Freedom-of-Information Files, Now Being Tested on Estonia's Records

Details

Maturity
Established
Promoter
University of Amsterdam (Informatics Institute) with the Dutch Government Organisation for Information Management (RvIHH)
Period
2022–2025
Region (NUTS)
NL32
Keywords
open government, freedom of information, natural language processing, digital archives

Context

The Netherlands' Open Government Act (Woo), which took effect in 2022, created a legal duty for public bodies to proactively publish far more of the documents they release under freedom-of-information requests, but there was no central way to find them, so the University of Amsterdam built Woogle, a free search engine that centralises Woo/Wob disclosures.

Objectives

Woogle set out to make more than three million scanned Dutch freedom-of-information dossiers searchable despite redactions, inconsistent labelling by releasing bodies, and bundled scans, and to test whether the same design could work beyond the Dutch context.

Activities

The pipeline applies optical character recognition to redacted documents, Page Stream Segmentation to reconstruct original document boundaries from bundled scans, and automated metadata extraction and classification to compensate for inconsistent labelling; in May 2025 this work was formalised into the ICAI OpenGov Lab, a joint research lab between the University of Amsterdam and the Dutch government's Information Management organisation (RvIHH), led by Dr. Maarten Marx and Dr. Jaap Kamps with Dr. David Graus as lab manager, extending the tooling with large-language-model and retrieval-augmented-generation research.

Results

Peer-reviewed documentation in the journal Data (MDPI) puts Woogle's holdings at more than three million FOIA dossiers, presented as over eleven million standardised pages, sourced from more than 1,000 distinct Dutch public bodies, and the same study reports that nearly 10,000 Estonian documents were successfully uploaded into the platform despite Estonia's ASiC-E and BDOC digital-signature formats and inconsistent, generic document titles creating real friction.

Conclusions

The published transfer test to Estonia is direct evidence that Woogle's approach is not purely a Dutch-language, Dutch-schema tool, though the Estonian extension remains a research-stage pilot of under 10,000 documents compared with the fully operational Dutch system, and no measurement of downstream citizen or journalist usage outcomes has yet been published.

Implementation

Indicative cost
Medium (€50k–€500k)
Time to results
Medium (1–3 years)
Staffing & skills
University of Amsterdam Informatics Institute researchers (Dr. Maarten Marx and Dr. Jaap Kamps as scientific directors), Dr. David Graus as ICAI OpenGov Lab manager, Dutch Government Organisation for Information Management (RvIHH) as government partner

Conditions for success

  • A legal mandate (the 2022 Woo Act) requiring proactive publication of FOI-released documents
  • A formal joint lab structure pairing university researchers with the government information-management body (RvIHH)
  • An OCR and document-classification pipeline robust enough to handle inconsistent metadata and labelling across 1,000+ releasing bodies

Common failure modes

  • Format and metadata inconsistency across releasing bodies, and across countries as seen with Estonia's ASiC-E/BDOC digital-signature formats and generic document titles, complicates automated processing
  • No published measurement yet of downstream usage outcomes such as journalist or citizen search success

Commonly funded by

Horizon Europe National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful