evidoria

← Back to browse

Good practice Imported

Assemblyline — Canada's Machine-Learning Malware Triage Platform Defending Government Networks

Canada · Ottawa · See the Canada profile · See the Ottawa profile

Evidence: Descriptive / self-reported Top 15% 80/100 · Ask Evidence Copilot about this practice

Canada's signals-intelligence agency runs Assemblyline, an open-sourced ML malware triage platform, plus a separate classifier flagging threats commodity antivirus misses across ~900,000 government devices, blocking 6.6 billion malicious actions daily.

1 billion+ files
Suspicious files scanned by Assemblyline (2023-2024 (CSE fiscal year))
308 (58 government, 250 critical infrastructure) organisations
Partner organisations using Assemblyline (government / critical infrastructure) (2023-2024)
35 %
Year-on-year growth in partner organisations (2023-2024 vs prior year)
~900000 devices
Devices covered by CSE sensors (2023-2024)
167 institutions
Federal institutions and Crown corporations covered (2023-2024)
6.6 billion actions/day
Average potentially malicious actions blocked per day (2023-2024 (up from 6.3 billion in 2022-23))
Assemblyline — Canada's Machine-Learning Malware Triage Platform Defending Government Networks

Details

Maturity
Established
Promoter
Communications Security Establishment (CSE) / Canadian Centre for Cyber Security
Period
2017-2024
Keywords
cybersecurity, malware detection, network defence, machine learning, CERT/CSIRT

Context

The Canadian Centre for Cyber Security, part of the Communications Security Establishment (CSE), operates sensors across Government of Canada networks and open-sourced Assemblyline in 2017, a machine-learning malware-analysis platform used to triage suspicious files at scale for government agencies and critical-infrastructure partners.

Objectives

Alongside Assemblyline, CSE data scientists built a separate machine-learning malware-classification model specifically aimed at catching custom, nation-state malware that commodity antivirus tools (built for generic threats) do not detect.

Activities

Flagged files from the ML classifier are quarantined, and the model is refined once commercial antivirus vendors catch up to the same threats, per CSE's own AI Strategy.

Results

CSE's Annual Report 2023-2024 states Assemblyline scanned over 1 billion suspicious files that year, with 308 partner organisations (58 government, 250 critical infrastructure) — up 35% year-on-year — and sensors covering roughly 900,000 devices across 167 federal institutions and Crown corporations. The Cyber Centre reports blocking an average of 6.6 billion potentially malicious actions per day, up from 6.3 billion in 2022-23. Independent reporting by the ICTC-CTIC policy institute and the Canadian magazine The Walrus corroborates the order of magnitude of blocked activity and sensor deployment, drawing on the same annual reports and an interview with CSE's chief.

Conclusions

CSE has not published a false-positive or accuracy rate specific to the ML classifier, nor a breakdown of how much blocked activity is specifically machine-learning-driven versus signature- or rule-based detection. No independent audit of detection accuracy exists; media corroboration confirms scale, not ML-specific performance claims.

Implementation

Indicative cost
High (€500k–€5M) — No budget figures are published in the source material; cost_band is a conservative estimate reflecting the scale of a national sensor network covering ~900,000 devices and 167 institutions, not a stated figure.
Time to results
Long (> 3 years) — Assemblyline open-sourced in 2017; continuously operated and scaled through at least the 2023-2024 reporting year.
Staffing & skills
CSE data scientists (built and maintain the ML malware-classification model), Canadian Centre for Cyber Security (operates network sensors, coordinates with government and critical-infrastructure partners)

Conditions for success

  • Open-sourcing Assemblyline in 2017 to enable broad adoption beyond CSE's own network (308 partner organisations by 2023-24)
  • Human-in-the-loop validation: ML-flagged files are quarantined and the model is refined once commercial antivirus vendors catch up

Common failure modes

  • No published false-positive or accuracy rate for the ML classifier creates an accountability/transparency gap
  • No independent audit of detection accuracy exists
  • No public breakdown separates ML-driven detections from signature- or rule-based detections within the reported blocked-action totals

Where it fits

Governance type
national signals-intelligence / cybersecurity agency
Scale
national (federal government plus critical-infrastructure partners)
Income level
high-income (non-EU, Canada)

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful