evidoria

← Back to browse

Good practice Imported

Gender Shades: The MIT Audit That Made Amazon, IBM and Microsoft Confront Facial-Recognition's Gender Gap

United States of America · Cambridge · See the United States of America profile · See the Cambridge profile

Evidence: Descriptive / self-reported Top 18% 89/100 · Ask Evidence Copilot about this practice

MIT's 2018 'Gender Shades' audit found facial-analysis systems misclassified darker-skinned women's gender up to 34.7% of the time vs under 1% for lighter-skinned men — evidence cited when Amazon paused police use of Rekognition and IBM exited facial recognition.

34.7 %
Gender-classification error rate, darker-skinned women (highest-error system) (2018)
<1 %
Gender-classification error rate, lighter-skinned men (2018)
31 %
Amazon Rekognition gender-misclassification rate, darker-skinned women (2019)
0 %
Amazon Rekognition gender-misclassification rate, lighter-skinned men (2019)

Details

Promoter
MIT Media Lab; Amazon; IBM
Period
2018-2020
Keywords
AI governance, facial recognition, policing, corporate policy

Context

In February 2018, MIT Media Lab researchers Joy Buolamwini and Timnit Gebru published "Gender Shades," a peer-reviewed audit (Proceedings of Machine Learning Research, FAT* 2018) of three commercial gender-classification systems: IBM, Microsoft and Face++.

Objectives

To quantify how accurately commercial facial-analysis systems classify gender across different skin tones, using a purpose-built, balanced benchmark dataset of parliamentarians from three African and three European countries.

Activities

The audit benchmarked IBM, Microsoft and Face++ on the balanced dataset. A 2019 follow-up audit by Deborah Raji and Buolamwini, "Actionable Auditing," specifically tested Amazon's Rekognition.

Results

All three systems performed far worse on darker-skinned women than on lighter-skinned men: error rates for darker-skinned women reached up to 34.7%, against error rates below 1% for lighter-skinned men. The 2019 follow-up found Amazon Rekognition misclassified darker-skinned women's gender roughly 31% of the time, versus 0% for lighter-skinned men. In June 2020, IBM announced it would discontinue general-purpose facial-recognition products, Amazon announced a one-year moratorium on police use of Rekognition (since extended indefinitely), and Microsoft paused sales of its facial-recognition technology to US police departments pending federal regulation.

Conclusions

Because the corporate policy changes coincided with mass protests over racial-justice issues broader than this specific audit, the research's causal share of the outcome cannot be cleanly isolated from the wider public pressure. The audit is nonetheless repeatedly and explicitly cited by the companies and independent press as a direct evidentiary trigger, and remains the most rigorously quantified account of gender-classification bias in commercial AI.

Implementation

Indicative cost
Low (< €50k) — An academic research audit; no implementation or programme costs are reported.
Time to results
Short (< 1 year) — Initial audit published February 2018; follow-up audit 2019; corporate responses (IBM's exit, Amazon's moratorium, Microsoft's pause) followed through 2020.
Staffing & skills
MIT Media Lab researchers (Joy Buolamwini, Timnit Gebru; 2019 follow-up with Deborah Raji)

Conditions for success

  • A purpose-built, phenotype- and gender-balanced benchmark dataset of parliamentarians from three African and three European countries
  • Independent peer review (FAT* 2018) that gave the findings credibility later cited directly by the audited vendors and by independent press

Common failure modes

  • Corporate policy responses coincided with the 2020 George Floyd protests, so the audit's own causal share of those changes cannot be cleanly isolated from broader public pressure

Commonly funded by

Philanthropic / foundation funding National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Replication kit

Reusable artefacts from this practice — as published by their sources.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful