AI Governance Institute
← Agentic AI
AGT · Agentic AIAGT-023High effortAgent-relevant

Agentic AI Security Assessment, CBRN and Cyber Espionage

Added June 2026

Assess AI agent deployments for high-consequence misuse, including chemical, biological, radiological, and nuclear facilitation or AI-orchestrated cyber espionage. Implement mitigations proportionate to identified risks.

Objective

Ensure organizations deploying agentic AI systems (AI that acts on its own) have explicitly assessed and mitigated the risk of those systems being weaponized or co-opted for high-consequence attacks. They should also be able to demonstrate to regulators and counterparties that this risk was considered and addressed.

Maturity Levels

1

Initial

Agentic AI systems are deployed without explicit assessment of high-consequence misuse risk. Security assessments cover conventional IT threats but not AI-specific attack methods.

2

Developing

Security teams are aware of CBRN and cyber espionage risk in AI contexts but assessments are informal and not documented. No structured threat model exists for agentic deployments.

3

Defined

A formal threat model assessment is conducted for each agentic AI deployment. It covers CBRN facilitation risk (based on the agent's tool access and the domains in which it can retrieve and synthesize information). It also covers cyber espionage risk (based on the agent's access to internal systems, credentials, and network). Mitigations are documented and implemented.

4

Managed

Assessments are reviewed annually and when an agent's capabilities or access scope changes materially. Results are reported to the AI governance committee and, for high-risk findings, to the board-level AI safety committee (BRD-003). Mitigations are tested.

5

Optimizing

External red-team exercises specifically targeting high-consequence misuse methods are conducted on a defined cadence for the highest-risk agentic systems. Assessment methodology is updated as the threat landscape evolves. The organization engages with AI safety research to stay current on emerging attack methods.

Get the free AI Governance Control Tracker

Get the free Excel tracker for all 132 governance controls. Score your maturity on Agentic AI Security Assessment, CBRN and Cyber Espionage and every other control, assign owners, and set deadlines.

  • 132 controls in Excel
  • Score maturity and assign owners
  • Track deadlines and regulation coverage

Includes AI Governance Weekly every Thursday. Unsubscribe anytime.

Evidence Requirements

What an auditor or assessor would expect to see for this control.

  • —Threat model assessment for each production agentic AI system covering CBRN facilitation and cyber espionage risk, with scoping rationale, findings, and implemented mitigations.
  • —Evidence of AI governance committee review for all assessments with material findings.
  • —Annual reassessment records or documented rationale for no material change.

Implementation Notes

Why this is a distinct control

General AI security controls address adversarial inputs (inputs crafted to trick the AI), prompt injection (hidden instructions that hijack the AI), and system compromise. This control addresses a different threat model (a structured view of who might cause harm and how). Here the adversary does not attack the agent itself but uses it to cause harm at a scale or in a domain that was not anticipated at deployment. The adversary might persuade the agent to help, exploit its access, or co-opt it as part of a larger attack chain.

This risk is not hypothetical. International AI safety research (including the 2026 International AI Safety Report and ARI's 2025 Safety Highlights) has documented cases of this. They include AI pulling together chemical, biological, radiological, and nuclear (CBRN) information, AI-orchestrated cyber espionage campaigns, and agentic AI systems used to execute cyberattacks with minimal human oversight.

Scoping the assessment

Not all agentic AI systems carry material CBRN or espionage risk. The assessment scope should be calibrated to the agent's capabilities:

CBRN facilitation risk indicators:

  • Agent has access to scientific literature retrieval (chemistry, biology, materials science).
  • Agent can synthesize multi-step research or planning outputs without human review.
  • Agent has access to procurement tools, supplier databases, or logistics systems.
  • Agent operates in life sciences, defense, or research contexts.

Cyber espionage risk indicators:

  • Agent has read access to internal systems, codebases, or credential stores (where passwords and access keys are kept).
  • Agent can connect to outside networks or send data to external systems.
  • Agent has access to communication systems (email, Slack, document repositories).
  • Agent has elevated privileges (broader system access than ordinary users) or can act under other identities.

For agents with no indicators in either category, a brief documented scoping justification is sufficient. Full assessment is required for agents with material indicators.

Assessment components

CBRN facilitation threat model:

  • List the domains of knowledge the agent can access and synthesize.
  • Assess whether retrieval and synthesis could meaningfully assist in CBRN attack planning.
  • Identify and implement safeguards, such as content filters on CBRN-adjacent queries and limits that block retrieval from relevant scientific domains. If the agent is a general-purpose large language model (LLM), use refusal training (teaching it to decline such requests).
  • Document residual risk (risk that remains after safeguards) and acceptance rationale.

Cyber espionage threat model:

  • Map all internal system access the agent holds.
  • Assess whether an adversary who controlled the agent (via prompt injection or direct access) could exfiltrate (secretly remove) sensitive data, use it to break into other systems, or scout systems for targets.
  • Implement mitigations: network egress controls (limits on what data can leave the network), data exfiltration monitoring, least-privilege access (only the access the agent needs), prompt injection defenses.
  • Conduct a red-team probe (a simulated attack by testers) if the risk level warrants it.

Regulatory and counterparty context

Several enterprise customers and government contractors now require explicit CBRN and dual-use (usable for both legitimate and harmful purposes) risk assessments for AI systems as a procurement condition. Defense, life sciences, and critical infrastructure organizations should expect these requirements to expand. Pre-emptive assessment positions the organization favorably.

Example Implementation

Agentic AI Security Assessment: Threat Model Summary

Agent: Internal Research Synthesis Agent | Assessment date: 2026-05-20 | Assessor: CISO + external red-team (Vendor X)

Scoping determination:

  • CBRN facilitation: Material risk indicators present. Agent has access to PubMed, bioRxiv, and internal research database. Can synthesize multi-step research summaries without per-output human review.
  • Cyber espionage: Moderate risk indicators. Agent has read access to internal Confluence and Jira. No credential store or network egress access.

CBRN facilitation findings:

  • The agent's retrieval scope includes biology, chemistry, and materials science literature without restriction.
  • Red-team probe: 15 adversarial queries designed to elicit synthesis of dual-use biological research were tested. 12 were refused by the underlying model. 3 produced partial outputs that stopped short of actionable synthesis.
  • Mitigation implemented: Retrieval scope restricted to exclude specific bioweapons-adjacent MeSH terms. Output classifier added that flags and routes for human review any synthesis touching designated dual-use categories.
  • Residual risk: Low. Accepted by CISO and CAIO on 2026-05-20.

Cyber espionage findings:

  • Agent can read Confluence pages including internal system architecture documents. Cannot write or exfiltrate via direct API.
  • Risk: An adversary controlling the agent via prompt injection could extract architecture documentation.
  • Mitigation implemented: Confluence retrieval scope limited to public and internal-general spaces. Architecture and security namespaces excluded from retrieval index.
  • Residual risk: Acceptable. Prompt injection defenses (AGT-002) further reduce risk.

Board committee reporting: No material unmitigated findings. Summary shared with AI Governance Committee on 2026-05-25. No escalation to Board AI Safety Committee required.

Control Details

Control ID
AGT-023
Typical owner
CISO / Chief AI Officer / Chief Risk Officer
Implementation effort
High effort
Agent-relevant
Yes

Tags

CBRNcyber espionagesecurity assessmenthigh-consequence riskagentic AIthreat modeling

Templates for this control