evidoria

← Back to browse

Good practice Imported

IGC and ZamStats Use GPT Models to Speed Up Zambia's Labour Force Survey Coding

Zambia · Lusaka · See the Zambia profile · See the Lusaka profile

Evidence: Quasi-experimental Top 25% 73/100 · Ask Evidence Copilot about this practice

Zambia's national statistics agency and the IGC tested seven LLMs against 1,059 human-coded Labour Force Survey responses: the best models beat enumerators by up to 17.8 percentage points on industry codes, at under $10 a year, projected to save over 170 working days annually.

1,059 responses
Benchmark sample size (2025)
11.8 percentage points
Occupation-code accuracy gain over enumerators (best models) (2025)
up to 17.8 percentage points
Industry-code accuracy gain over enumerators (best models) (2025)
60-63 %
Best-model occupation-code exact-match accuracy (2025)
57-66 %
Best-model industry-code exact-match accuracy (2025)
<10 USD
Cost to classify a full annual LFS sample (GPT-5 Nano) (2025)
~43 working days/year
Projected annual time saved, ZamStats headquarters staff (projected)
~130 working days/year
Projected annual time saved, field enumerators (projected)
IGC and ZamStats Use GPT Models to Speed Up Zambia's Labour Force Survey Coding IGC and ZamStats Use GPT Models to Speed Up Zambia's Labour Force Survey Coding

Details

Maturity
Pilot
Promoter
Zambia Statistics Agency (ZamStats) & International Growth Centre (IGC)
Period
2025
Keywords
official statistics, labour market, public administration

Context

Zambia's quarterly Labour Force Survey (LFS), run by the Zambia Statistics Agency (ZamStats) with the Ministry of Labour and Social Security, interviews 10,400 households a year. Enumerators assign each respondent a four-digit occupation code (436 possible ISCO codes) and industry code (419 possible ISIC codes) from verbal descriptions collected in the field, a process the survey team describes as time-consuming and prone to error.

Objectives

In 2025, ZamStats partnered with the International Growth Centre (IGC) and the University of Oxford to test whether large language models could code these responses better and faster than human enumerators.

Activities

Researchers built a retrieval-augmented-generation pipeline with few-shot prompting and benchmarked seven LLMs, including GPT-4 Turbo, GPT-5 mini and GPT-5 Nano, against 1,059 LFS responses that ZamStats headquarters staff had independently coded as ground truth. ZamStats and IGC also built an open, no-code interface for the pipeline, targeting full implementation by the end of 2025.

Results

Every LLM tested outperformed human enumerators by a statistically significant margin (p<0.001). The best-performing models reached 60-63% exact-match accuracy on occupation codes (11.8 percentage points above enumerators) and 57-66% on industry codes (up to 17.8 percentage points above enumerators). Classifying an entire annual LFS sample costs under $10 with GPT-5 Nano, and the team calculates the approach could save roughly 43 working days a year for headquarters staff and 130 working days a year for field enumerators, using conservative time estimates.

Conclusions

The authors note real limitations: the 1,059-response sample is too small to compare LLMs against each other with statistical confidence, the enumerator 'ground truth' itself contains coding errors, and results are demonstrated only for Zambia's LFS so far, though the team argues the method could scale to Zambia's four-million-household census and to labour surveys in other countries.

Implementation

Indicative cost
Low (< €50k) — Classifying an entire annual LFS sample costs under $10 using GPT-5 Nano.
Time to results
Short (< 1 year) — Benchmarking conducted in 2025 with full implementation targeted for end of 2025; the team argues the method could scale to Zambia's four-million-household census and to other countries' labour surveys.
Staffing & skills
Zambia Statistics Agency (ZamStats) headquarters staff, Ministry of Labour and Social Security, International Growth Centre (IGC) and University of Oxford researchers, Field enumerators who collect the underlying verbal descriptions

Conditions for success

  • An independently ground-truthed benchmark dataset to validate model performance before rollout
  • A retrieval-augmented-generation pipeline with few-shot prompting tailored to the coding taxonomy (ISCO/ISIC)
  • An open, no-code interface so ZamStats staff can operate the pipeline without programming skills

Common failure modes

  • The 1,059-response benchmark sample is too small to compare LLMs against each other with statistical confidence.
  • The enumerator 'ground truth' itself contains coding errors.
  • Results are demonstrated only for Zambia's LFS so far, not yet other surveys or countries.

Where it fits

Governance type
national statistics agency, with an external research partnership
Scale
national statistical survey (10,400 households/year)
Income level
lower-middle-income

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Replication kit

Reusable artefacts from this practice — as published by their sources.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful