Agent Prompt Injection Defense
Added May 2026
Protect AI agents from prompt injection attacks, hidden instructions planted in outside content that take over agent behavior.
Objective
Prevent agents from being redirected by malicious instructions in tool outputs (results from software the agent calls), user-supplied content, or data pulled from the web.
Maturity Levels
Initial
No prompt injection defenses exist; agents process external content without filtering.
Developing
Engineering teams are aware of prompt injection but defenses are inconsistent and undocumented.
Defined
Input sanitization and content-boundary enforcement are applied to all agent inputs from external sources.
Managed
Red team testing for prompt injection is conducted quarterly; findings are tracked to remediation.
Optimizing
Automated injection attempt detection feeds continuous improvement of agent system prompts and input filters.
Evidence Requirements
What an auditor or assessor would expect to see for this control.
- —Input sanitization (screening of incoming content) configuration documentation listing active filters and the rules that keep external content separate from instructions
- —Quarterly red team testing (staged attacks by testers) results: number of injection attempts tested, pass/fail outcome, and remediation actions for any failures
- —Injection attempt detection logs reviewed on a defined cadence, with escalation records for confirmed attempts
- —System prompt (the agent's standing instructions) documentation demonstrating structural separation of agent instructions from external data
- —Post-remediation re-test results confirming findings from red team exercises were resolved
Implementation Notes
Key steps
- Treat all external content (web pages, emails, documents, API responses from other software systems) as untrusted, enforce a strict boundary between agent instructions and external data.
- Use structural separation: pass external content in clearly labeled data fields rather than mixing it into the agent's instructions.
- Test agents explicitly for indirect injection (instructions hidden in content the agent reads): submit tool results containing instruction-like text and verify the agent does not follow them.
- Monitor for unusual agent actions that stray from the original task, as injection attacks often push agents beyond what they were asked to do.
Example Implementation
Sales team using an AI agent to process and route inbound email inquiries
Prompt Injection Defense Policy: Email Intake Agent
Content boundary rule: All email body content is passed inside structured <email> tags. The system prompt explicitly instructs the agent to treat any content inside <email> as untrusted data and to ignore instructions found within it. External content is never appended to the instruction context.
Permitted action set: {classify, route, flag_for_review}. Any agent response outside this set is logged as a potential injection attempt and held for human review.
Input filters (applied before agent receives content):
- Strip strings matching: "ignore previous instructions", "you are now", "new task:", "system:", "[INST]"
- Flag any content claiming to originate from the system, operator, or a higher-trust agent
Monitoring: Injection attempt log reviewed weekly; patterns feed updates to system prompt and filter rules
Red team schedule: Quarterly, 20 emails with embedded instruction-like payloads; pass criterion: 0 instructions followed
