Question 20 of 53
Is our AI red-teaming rigorous enough?
By Cody Maxwell · AI Governance Institute · February 2026 · Updated September 2026
Set pass/fail criteria for testing high-risk AI before deployment. Cover toxicity, data leakage, jailbreaking, and misuse.
If you only do 3 things, do this:
- 1.Define pass/fail criteria before red-teaming begins, not after. What is an unacceptable failure for your specific use case? A content moderation system and a loan decision system have very different tolerances.
- 2.Re-test after every remediation. Fixes introduce their own failure modes. A passed re-test is a prerequisite for deployment, not a formality.
- 3.Document all findings, including attacks that didn't succeed. A red-team log that only shows failures creates a false picture of safety.
The Situation
Who this is for: Security, machine learning engineering, and compliance teams responsible for pre-deployment testing of high-risk AI systems
When you need this: Before deploying any high-risk AI system, after major model updates, and on a scheduled basis for production generative AI
The Decision
Does our AI system have failure modes that would be unacceptable in production, and have we tested rigorously enough to be confident we've found the important ones?
The Steps
- 1Define pass/fail criteria for the specific system and use case: be explicit about what constitutes an unacceptable outcome
- 2Assemble a red team with the mandate to actively try to break the system (do not settle for a checklist review)
- 3For generative AI: test attempts to talk it past its safety rules (jailbreaking), hidden instructions that hijack it (prompt injection), attempts to pull out confidential data, and misuse scenarios specific to your deployment
- 4For decision-making AI: test borderline cases where a small change flips the decision, probe for bias, and check how it handles unusual inputs unlike anything it was trained on
- 5Document all findings: attack description, result, severity, and recommended remediation
- 6Remediate findings above the acceptable threshold; re-test to verify fixes and confirm no new failures were introduced
- 7Archive red-team findings and re-test results as part of the system's deployment documentation
The Artifacts
- —Red-team scope definition template (system description, failure definition, pass/fail criteria)
- —Attack scenario library (jailbreaking, prompt injection, data extraction, bias probes, robustness tests)
- —Red-team findings log template (attack, result, severity, remediation, re-test status)
- —Pre-deployment security checklist for AI systems
- —Red-team report template (executive summary plus technical findings)
The Output
A documented red-team assessment for every high-risk AI system, with pass/fail criteria met, all high-severity findings remediated and re-tested, and results archived as deployment documentation.
Red-teaming is not optional for high-risk AI
Red-teaming, adversarial testing designed to find failure modes before deployment, is increasingly expected by regulators and required by internal governance standards for high-risk AI systems. The EU AI Act requires accuracy, robustness, and cybersecurity testing for high-risk systems. NIST's AI RMF and the Generative AI Profile both address adversarial testing as a component of responsible deployment.
A red-team exercise that does not find anything is usually an exercise that did not try hard enough. Effective red-teaming requires people who are actively trying to make the system fail, with the creativity and persistence to find non-obvious failure modes rather than just checking a predetermined list.
What red-teaming should cover
For generative AI systems, red-teaming should cover: jailbreaking attempts designed to bypass safety controls and produce prohibited content; prompt injection attacks that attempt to override system instructions through user input; data extraction probes designed to elicit training data or confidential system information; and misuse scenarios specific to the deployment context, such as generating fraudulent documents or manipulating decisions.
For decision-making AI systems, testing should cover: borderline cases where a small change in the input flips the decision; inputs designed to exploit known model weaknesses or biases; and how the system holds up when real-world data shifts away from its training data or when it sees inputs unlike anything it was trained on.
Pass/fail criteria and documentation
Define pass/fail criteria before red-teaming begins, not after. What constitutes an unacceptable failure for your specific use case? A content moderation system and a loan decision system have very different failure tolerances. Criteria should be specific, measurable, and tied to the risk level and regulatory requirements of the system.
Document all red-team findings, including the attacks attempted, the results, the severity assessment, and the remediation taken. Findings that do not result in immediate remediation should be tracked as known risks with owners and timelines. Re-test after remediation to verify that fixes are effective and have not introduced new failure modes. Red-team documentation becomes part of the system's audit trail and may be reviewed by regulators or in litigation.
Turn this guidance into an implementation plan
Get the free Excel tracker for all 132 governance controls. Score maturity, assign owners, and set deadlines, including this playbook's 19 related controls.
- 132 controls in Excel
- Score maturity and assign owners
- Track deadlines and regulation coverage
Includes AI Governance Weekly every Thursday. Unsubscribe anytime.
Governance Controls
Operational controls that implement the guidance in this playbook.
Recent Coverage
News and developments relevant to this playbook topic.
