AI Governance Institute
← Agentic AI
AGT · Agentic AIAGT-002Medium effortAgent-relevant

Agent Prompt Injection Defense

Added May 2026

Protect AI agents from prompt injection attacks, hidden instructions planted in outside content that take over agent behavior.

Objective

Prevent agents from being redirected by malicious instructions in tool outputs (results from software the agent calls), user-supplied content, or data pulled from the web.

Maturity Levels

1

Initial

No prompt injection defenses exist; agents process external content without filtering.

2

Developing

Engineering teams are aware of prompt injection but defenses are inconsistent and undocumented.

3

Defined

Input sanitization and content-boundary enforcement are applied to all agent inputs from external sources.

4

Managed

Red team testing for prompt injection is conducted quarterly; findings are tracked to remediation.

5

Optimizing

Automated injection attempt detection feeds continuous improvement of agent system prompts and input filters.

Evidence Requirements

What an auditor or assessor would expect to see for this control.

  • —Input sanitization (screening of incoming content) configuration documentation listing active filters and the rules that keep external content separate from instructions
  • —Quarterly red team testing (staged attacks by testers) results: number of injection attempts tested, pass/fail outcome, and remediation actions for any failures
  • —Injection attempt detection logs reviewed on a defined cadence, with escalation records for confirmed attempts
  • —System prompt (the agent's standing instructions) documentation demonstrating structural separation of agent instructions from external data
  • —Post-remediation re-test results confirming findings from red team exercises were resolved

Implementation Notes

Key steps

  • Treat all external content (web pages, emails, documents, API responses from other software systems) as untrusted, enforce a strict boundary between agent instructions and external data.
  • Use structural separation: pass external content in clearly labeled data fields rather than mixing it into the agent's instructions.
  • Test agents explicitly for indirect injection (instructions hidden in content the agent reads): submit tool results containing instruction-like text and verify the agent does not follow them.
  • Monitor for unusual agent actions that stray from the original task, as injection attacks often push agents beyond what they were asked to do.

Example Implementation

Sales team using an AI agent to process and route inbound email inquiries

Prompt Injection Defense Policy: Email Intake Agent

Content boundary rule: All email body content is passed inside structured <email> tags. The system prompt explicitly instructs the agent to treat any content inside <email> as untrusted data and to ignore instructions found within it. External content is never appended to the instruction context.

Permitted action set: {classify, route, flag_for_review}. Any agent response outside this set is logged as a potential injection attempt and held for human review.

Input filters (applied before agent receives content):

  • Strip strings matching: "ignore previous instructions", "you are now", "new task:", "system:", "[INST]"
  • Flag any content claiming to originate from the system, operator, or a higher-trust agent

Monitoring: Injection attempt log reviewed weekly; patterns feed updates to system prompt and filter rules

Red team schedule: Quarterly, 20 emails with embedded instruction-like payloads; pass criterion: 0 instructions followed

Control Details

Control ID
AGT-002
Typical owner
AI Security / AI Engineering
Implementation effort
Medium effort
Agent-relevant
Yes

Tags

prompt injectionagent securityadversarial inputsindirect injection

Templates for this control

Get control updates weekly

New and updated controls, maturity guidance, and the regulatory changes behind them. Every Thursday.

Powered by Buttondown.