evidoria

← Back to browse

Good practice Imported

CISA's AI-Enabled Vulnerability Detection Pilot — Testing Whether AI Tools Actually Improve Federal Cyber Threat Hunting

United States of America · Arlington · See the United States of America profile

Evidence: Quasi-experimental Top 35% 67/100 · Ask Evidence Copilot about this practice

In a 2023–24 operational pilot ordered under a presidential AI directive, CISA compared AI/LLM-based vulnerability-detection tools against its existing methods on real federal networks — and reported the gains were often negligible, not a clear win for AI.

Details

Maturity
Pilot
Promoter
Cybersecurity and Infrastructure Security Agency (CISA)
Period
2023–2024
Keywords
cybersecurity, critical infrastructure, vulnerability management, homeland security

Context

The US Cybersecurity and Infrastructure Security Agency (CISA) ran an operational pilot from late 2023 to early 2024 to test whether commercially available AI and large language model (LLM) based vulnerability-detection tools outperform CISA's existing, non-AI detection methods. The pilot combined security assessments of real federal partner networks with tests inside a controlled environment, evaluating AI products available as of December 31, 2023, with priority given to newer LLM-based tools.

Objectives

The pilot was ordered under a presidential directive on AI and critical-infrastructure security, requiring CISA to deliver its findings to the Department of Homeland Security.

Activities

CISA delivered its findings to the Department of Homeland Security on July 26, 2024.

Results

The report concluded that AI currently works best as a supplement to existing vulnerability-detection tools rather than a replacement for them: in some cases, the time analysts needed to learn the new AI capabilities was substantial while the resulting improvement in detection was negligible, and AI tools could behave unpredictably in ways that were difficult to troubleshoot.

Conclusions

CISA said it would continue monitoring and testing AI vulnerability-detection products as the market matures, rather than adopting them at scale on the strength of this pilot. The findings are a rare example of a government AI evaluation publishing an inconclusive, cautionary result rather than a success narrative.

Implementation

Indicative cost
Low (< €50k) — Operational pilot conducted with existing agency resources and commercially available AI products; no separate budget figure published.
Time to results
Short (< 1 year) — Ran from late 2023 to early 2024; findings delivered to DHS on 26 July 2024.
Staffing & skills
Cybersecurity and Infrastructure Security Agency (CISA), Federal partner network operators, Department of Homeland Security (recipient of findings)

Conditions for success

  • Testing AI tools against an existing non-AI baseline on real production networks, not only synthetic tests
  • Willingness to publish an inconclusive/negative finding rather than only favourable results

Common failure modes

  • Time analysts needed to learn new AI tools was substantial while detection improvement was often negligible
  • AI tools could behave unpredictably in ways difficult to troubleshoot

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful