In a 2023–24 operational pilot ordered under a presidential AI directive, CISA compared AI/LLM-based vulnerability-detection tools against its existing methods on real federal networks — and reported the gains were often negligible, not a clear win for AI.
Details
Maturity
Pilot
Promoter
Cybersecurity and Infrastructure Security Agency (CISA)
The US Cybersecurity and Infrastructure Security Agency (CISA) ran an operational pilot from late 2023 to early 2024 to test whether commercially available AI and large language model (LLM) based vulnerability-detection tools outperform CISA's existing, non-AI detection methods. The pilot combined security assessments of real federal partner networks with tests inside a controlled environment, evaluating AI products available as of December 31, 2023, with priority given to newer LLM-based tools.
Objectives
The pilot was ordered under a presidential directive on AI and critical-infrastructure security, requiring CISA to deliver its findings to the Department of Homeland Security.
Activities
CISA delivered its findings to the Department of Homeland Security on July 26, 2024.
Results
The report concluded that AI currently works best as a supplement to existing vulnerability-detection tools rather than a replacement for them: in some cases, the time analysts needed to learn the new AI capabilities was substantial while the resulting improvement in detection was negligible, and AI tools could behave unpredictably in ways that were difficult to troubleshoot.
Conclusions
CISA said it would continue monitoring and testing AI vulnerability-detection products as the market matures, rather than adopting them at scale on the strength of this pilot. The findings are a rare example of a government AI evaluation publishing an inconclusive, cautionary result rather than a success narrative.
Implementation
Indicative cost
Low (< €50k) — Operational pilot conducted with existing agency resources and commercially available AI products; no separate budget figure published.
Time to results
Short (< 1 year) — Ran from late 2023 to early 2024; findings delivered to DHS on 26 July 2024.
Staffing & skills
Cybersecurity and Infrastructure Security Agency (CISA), Federal partner network operators, Department of Homeland Security (recipient of findings)
Conditions for success
Testing AI tools against an existing non-AI baseline on real production networks, not only synthetic tests
Willingness to publish an inconclusive/negative finding rather than only favourable results
Common failure modes
Time analysts needed to learn new AI tools was substantial while detection improvement was often negligible
AI tools could behave unpredictably in ways difficult to troubleshoot
Commonly funded by
National / regional programmes
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Do you run this practice?
Claim it —
verified implementers get a public contact pathway and can propose corrections.
Data sources
Where this practice's information was retrieved from, and when.
An open-source Analyzer module, built by STACC for Estonia's Information System Authority (RIA), applies adaptive statistical models to X-Road's data-exchange …
Since 2020, Czech Republic's NÚKIB has run Flowmon's machine-learning network anomaly detection—combining ML, heuristics and behavioural analytics across ~40 algorithms—across …