evidoria

← Back to browse

Good practice Imported

AutoML Dropout-Risk Model for Tanzania's Secondary Schools — University of Dar es Salaam

Tanzania · Dar es Salaam · See the Tanzania profile · See the Dar es Salaam profile

Evidence: Observational / pre–post Top 89% 27/100 · Ask Evidence Copilot about this practice

University of Dar es Salaam researchers built an automated machine-learning model to flag secondary-school students at risk of dropping out, using Tanzania's national Twaweza Uwezo household-assessment data as dropout rates rose from 3.8% (2018) to 4.2% (2019).

3.8 %
National secondary-school dropout rate (2018)
4.2 %
National secondary-school dropout rate (2019)
99.8 %
Decision Tree classifier accuracy
99.6 %
K-Nearest Neighbours classifier accuracy
99 %
Multi-Layer Perceptron classifier accuracy
97 %
Naive Bayes classifier accuracy

Details

Promoter
University of Dar es Salaam, Department of Computer Science and Engineering
Period
2018–2022
Keywords
computer science, machine learning, secondary education, dropout prevention

Context

Tanzania's secondary schools have struggled with a persistent and slightly worsening dropout problem, with the national rate rising from 3.8% in 2018 to 4.2% in 2019, and with limited formal tools available to identify the root causes or project which students are most at risk.

Activities

Researchers Yuda N. Mnyawami, Hellen H. Maziku and Joseph C. Mushi, in the Department of Computer Science and Engineering at the University of Dar es Salaam, responded with an Automated Machine Learning (AutoML) approach that selects the hyperparameters, features and algorithm best suited to the available data, rather than relying on a single hand-tuned model. They trained and tested the approach on Tanzania's Twaweza Uwezo household learning-assessment dataset, which covers households across the country, comparing four classifiers.

Results

The reported accuracy was high across all four models — Decision Tree 99.8%, K-Nearest Neighbours 99.6%, Multi-Layer Perceptron 99% and Naive Bayes 97% — with the authors arguing that automated hyperparameter and feature selection outperformed the manually tuned models used in earlier dropout-prediction studies.

Conclusions

These figures should be read cautiously: accuracy this high is often a sign of a small or imbalanced evaluation set rather than a production-ready classifier, and the paper does not report validation against an independent Ministry of Education dataset or a real classroom pilot. A related follow-up by overlapping authors extends a similar Bayesian-optimisation approach to secondary schools elsewhere in Sub-Saharan Africa, suggesting this is part of an active regional research programme — but as of publication, no ministry has deployed the model as an operational early-warning tool, and no intervention outcomes have been measured.

Implementation

Indicative cost
Low (< €50k)
Time to results
Short (< 1 year)
Staffing & skills
Yuda N. Mnyawami, Hellen H. Maziku, Joseph C. Mushi — Department of Computer Science and Engineering, University of Dar es Salaam

Conditions for success

  • used an existing national household dataset (Twaweza Uwezo) rather than requiring new data collection
  • compared four classifiers with automated hyperparameter/feature selection rather than a single hand-tuned model

Common failure modes

  • unusually high accuracy (up to 99.8%) is a plausible sign of overfitting or a small/imbalanced evaluation set rather than production-ready performance
  • no validation reported against an independent Ministry of Education dataset
  • no real classroom pilot or deployed early-warning tool exists; no ministry has adopted the model operationally
  • no intervention outcomes have been measured, since the model has not been deployed

Where it fits

Governance type
academic research at a public university
Scale
national dataset, pre-deployment research stage
Income level
low-income country (Tanzania)

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful