evidoria

← Back to browse

Good practice Imported

NASA's AI-Generated Metadata Tags — Machine-Learning Auto-Tagging for Federal Open Data Discovery

United States of America · Washington, D.C. · See the United States of America profile · See the Washington, D.C. profile

Evidence: Descriptive / self-reported Top 12% 80/100 · Ask Evidence Copilot about this practice

Since around 2019, NASA's Scientific and Technical Information Program has used a machine-learning model — trained on 3.5 million tagged documents — to auto-tag NASA's open data against a taxonomy of roughly 20,000 standardized keywords, and open-sourced the tooling.

Details

Maturity
Established
Promoter
NASA Scientific and Technical Information (STI) Program / NASA OCIO Data Analytics Team
Period
2018–present
Keywords
open data, machine learning, information management, federal government

Context

NASA's Scientific and Technical Information (STI) corpus spans scientific articles, technical reports, software repositories and other unstructured content scattered across multiple systems, historically tagged by hand with roughly ten keywords per document drawn from an evolving taxonomy of more than 20,000 standardised terms.

Objectives

The STI team and NASA's OCIO Data Analytics Team set out to build a machine-learning and NLP-based automated tagging system to reduce the manual burden of tagging and improve discoverability of NASA's technical content.

Activities

The system was trained on the NASA Technical Reports Server's corpus of 3.5 million already-tagged documents, then released as open-source software (the concept-tagging-training and concept-tagging-api repositories on GitHub) and documented as a Federal Data Strategy 'proof point' case study, published May 2019. Its use has since expanded beyond original article-tagging to other unstructured NASA data.

Results

No public accuracy, usage or before/after discoverability figures have been published, so the tool's measured impact on data findability remains undemonstrated even though the system has been operating for several years.

Conclusions

The case study explicitly frames the tool as a template other federal agencies can replicate on their own text corpora, backed by open-source code rather than a proprietary black box.

Implementation

Indicative cost
Low (< €50k) — Built in-house by existing NASA STI Program and OCIO staff using an existing tagged corpus; no separate procurement or licensing cost disclosed; code released free as open source.
Time to results
Medium (1–3 years) — System built and documented by 2018–2019 (Federal Data Strategy proof point published May 2019); operating and expanding in scope through 2026 (present).
Staffing & skills
NASA Scientific and Technical Information (STI) Program, NASA OCIO Data Analytics Team

Conditions for success

  • A large, already human-tagged corpus (3.5 million documents) available to train the model
  • An existing, standardised keyword taxonomy (20,000+ terms) to tag against
  • Releasing the code as open source so other agencies can reuse it directly

Common failure modes

  • No public accuracy, usage or before/after discoverability metrics have been published, leaving the tool's real-world impact unverified
  • Tagging quality depends on the historical taxonomy staying current as NASA's technical content evolves

Where it fits

Governance type
US federal agency (NASA)
Scale
national/federal
Income level
high-income

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Replication kit

Reusable artefacts from this practice — as published by their sources.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful