AI Governance Institute
← News
Research2026-09-16

Multimodal Prompt Injection Exposes Structural Gap in Agent Red-Teaming

What happened

Security research organization Co-RE published the AI Agent Attack Techniques & Countermeasures catalog in September 2026. The document systematically maps how adversaries can inject malicious instructions into AI agents not only through direct text prompts but also through images, documents, audio files, and combinations of those modalities. It specifically highlights zero-click and hidden injection paths that require no user interaction to trigger and can cause agents to disclose sensitive data or take actions outside their sanctioned scope. The catalog arrives as a series of documented incidents has shown real-world exploitation of these exact channels. Prior research, including findings on zero-click prompt injection escaping coding agent sandboxes and hidden HTML prompt injection achieving 100% success rates against email summarizers, confirm that these are active, not theoretical, attack classes. The catalog gives compliance and security teams a structured reference for evaluating whether their current test programs cover the full threat surface.

Why it matters

  • ·Most enterprise red-teaming programs were designed for text-based model testing. Agents that process images, audio, or documents face injection paths those programs do not evaluate, leaving a documented and exploitable gap in pre-deployment assurance.
  • ·Regulators and auditors increasingly expect red-teaming scope to match deployment reality. A program that misses multimodal injection channels may not satisfy conformity assessment requirements under frameworks like the EU AI Act or voluntary commitments tied to Five Eyes agentic AI security guidance.
  • ·Zero-click injection paths, where no user action is needed to trigger an agent into unauthorized behavior, create direct exposure for data-handling, audit-trail, and incident-response controls. A successful hidden injection that causes data exfiltration may not appear in standard agent logs, compounding the governance failure.

Governance controls affected

What to do now

  • ☐Audit your current AI agent red-teaming scope and confirm it explicitly covers image, document, audio, and multimodal input channels, not only direct text prompts.
  • ☐Review deployment configurations for any agent that ingests user-supplied files or external media and verify that untrusted content is treated as a potential injection vector before processing.
  • ☐Update your AGT-016 (Agentic AI Deployment Readiness Assessment) checklist to require documented multimodal injection testing as a gate condition before agent deployment.
  • ☐Confirm that agent action audit trails capture actions triggered by non-text inputs, including file-processing events, so that hidden injection attempts are recoverable for incident review.
  • ☐Brief your AI red-teaming or penetration testing vendor on the Co-RE catalog and require that their next engagement addresses zero-click and hidden injection paths explicitly.

What to watch next

The pace of documented multimodal injection incidents is accelerating, and at least two regulatory signals suggest that structured red-teaming scope will become a formal requirement rather than a best practice. The Five Eyes agentic AI security guidance already treats sandboxing and adversarial testing as baseline expectations. Compliance teams should also watch for updates to NIST and CSA agent security guidance, both of which are active workstreams that may formalize multimodal injection testing requirements. Any enterprise that has not updated its red-teaming protocol since deploying multimodal-capable agents should treat that gap as a pre-audit remediation priority.

Related Coverage

Research2026-10-03

Orchestration Framework Flaws Make AI Workflow Pipelines a Primary Attack Target

Research published by Help Net Security finds that agent orchestration frameworks including Flowise and Langflow are among the most actively targeted systems in current vulnerability disclosures. Attackers use prompt injection and manipulated workflow configuration files to reach code execution points inside enterprise AI pipelines. Organizations running agentic workflows need isolation, configuration validation, and red-team coverage at the orchestration layer, not just at the model level.

Research2026-10-02

Six Agentic Failure Modes Show Soft Guardrails Are Not Enough

A practitioner analysis published by CSO Online identifies six named failure modes in deployed AI agents, including prompt injection, context manipulation, and authorization abuse. The analysis draws on real incidents, including the OpenAI Atlas browser hijack and the Microsoft 365 Copilot EchoLeak exploit. It concludes that enterprises relying solely on vendor-configured content filters and system-prompt instructions have not closed the control loop.

Research2026-09-30

OpenAI's GPT-5.6 Red-Team Finds Self-Replicating Prompt Injection

OpenAI disclosed in September 2026 that its GPT-5.6 model is susceptible to self-replicating prompt injection attacks, discovered during internal red-teaming by an automated agent called GPT-Red. The attacks spread malicious instructions across connected systems such as email and calendars without human interaction. No exploitation outside testing environments was confirmed, but OpenAI is now using the attack patterns in model training.