AI Governance Institute
← All news

Adversarial Jailbreak

An adversarial jailbreak is an attack where someone feeds a generative AI system specially crafted prompts or inputs designed to bypass its safety guidelines and make it produce harmful content it was trained to refuse. These attacks matter for governance because they reveal gaps between a system's stated values and its actual behavior, exposing your organization to reputational risk, regulatory scrutiny, and potential liability if a jailbroken AI generates discriminatory, illegal, or dangerous outputs under your control. Enterprise teams need to test for and monitor these vulnerabilities as part of their AI risk assessment and incident response planning.

1 item