evidoria

← Back to browse

Good practice Imported

Harmonizing 4.1 Million Multilingual Tax Records — Rwanda Revenue Authority's NLP Clustering Pipeline

Rwanda · Kigali · See the Rwanda profile · See the Kigali profile

Evidence: Descriptive / self-reported Top 50% 60/100 · Ask Evidence Copilot about this practice

A study built with Rwanda Revenue Authority data used MiniLM, UMAP and HDBSCAN to harmonize 4.1 million multilingual product names from EBM tax records into 425,103 standardized names, exposing pricing anomalies.

4,124,005
EBM and customs records processed (FY2020-2022)
425,103
Standardized product names produced
6,295
Fine-grained HDBSCAN clusters produced
167,085
Records flagged as noise (typos/rare labels)
73.8 %
Non-English product-name records in the dataset

Details

Maturity
Pilot
Promoter
Rwanda Revenue Authority (Strategy and Risk Analysis Department), research by Adventist University of Central Africa
Period
2020–2025
Keywords
tax administration, customs, fraud detection

Context

Since 2013 the Rwanda Revenue Authority (RRA) has digitized invoicing through Electronic Billing Machines (EBMs), reaching structured, item-level data with EBM v2.0/2.1. But manual, multilingual entry of product names — 'sugar' recorded interchangeably as sucre, isukari or sukari — left records inconsistent and hard to analyse for fraud.

Objectives

Harmonize multilingual, inconsistently entered product names into standardized categories so that identical products can be compared across languages for pricing and fraud analysis.

Activities

A 2025 study conducted with the RRA's Strategy and Risk Analysis Department, which supplied datasets and technical input, built a pipeline of text cleaning, language detection and translation, MiniLM sentence embeddings, PCA/UMAP dimensionality reduction, and KMeans plus HDBSCAN clustering, processing 4,124,005 EBM and customs (Electronic Single Window) records from FY2020-2022, of which 73.8% were non-English (Kinyarwanda, French or Swahili).

Results

The pipeline produced 425,103 standardized product names, 20 broad KMeans categories and 6,295 fine-grained HDBSCAN clusters (plus 167,085 flagged noise points such as typos and rare labels). Grouping identical products across languages let researchers compare import/purchase values against sales prices for the same item, surfacing real pricing discrepancies — for example one 'a3 paper' cluster ranging from 1,000 to 15,000 Rwandan francs, consistent with underpricing or misreporting.

Conclusions

This is a peer-reviewed research study (an MSc-level thesis project) built on real RRA administrative data rather than a confirmed live production deployment. The authors describe the pipeline's modular design as scalable to other sectors and languages within government data systems.

Implementation

Indicative cost
Low (< €50k)
Time to results
Short (< 1 year) — Processed FY2020-2022 administrative records; published as a peer-reviewed journal article in 2025.
Staffing & skills
Rwanda Revenue Authority Strategy and Risk Analysis Department (data and technical input), Researcher(s) at Adventist University of Central Africa (MSc thesis project)

Conditions for success

  • Access to structured, item-level EBM v2.0/2.1 invoicing data
  • Language detection/translation capability to handle Kinyarwanda, French, Swahili and English product names
  • Modular pipeline design enabling reuse across other product categories and languages

Common failure modes

  • 167,085 records (about 4%) remained unresolved as noise points (typos, rare labels), requiring further cleaning
  • Not yet confirmed as a live production fraud-detection workflow inside RRA

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Replication kit

Reusable artefacts from this practice — as published by their sources.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful