Question 50 of 52
How do we audit our AI governance program, not just individual AI systems?
By Cody Maxwell · AI Governance Institute · August 2026
A method for auditing the AI governance program itself against maturity levels, separate from auditing any single AI system for compliance — what an internal audit function should test, and what evidence counts.
If you only do 3 things, do this:
- 1.Program auditing and system auditing are different exercises. A system audit checks one deployment; a program audit checks whether your governance function works at all, across every system.
- 2.Score maturity per control domain, not with one overall grade. A program can be mature in documentation and immature in monitoring at the same time, and a single score hides that.
- 3.Treat "the control is documented" and "the control is enforced" as two separate audit findings. The gap between them is where most governance programs actually fail.
The Situation
Who this is for: Internal audit, compliance leads, or risk committees needing to assess whether the AI governance program itself is functioning, not just whether one system passed review
When you need this: On a recurring audit cycle, before a board or regulator asks for assurance over the whole program, or after an incident raises doubt about whether existing controls are real
The Decision
Is the AI governance program operating at the maturity level the organization believes it is, and where specifically does the evidence fall short?
The Steps
- 1Define the control domains in scope for the audit (for example: inventory, risk classification, human oversight, monitoring, incident response, documentation, vendor risk)
- 2For each domain, define what maturity level 1 through 4 or 5 looks like in concrete, testable terms rather than descriptive language
- 3Sample a cross-section of AI systems from the inventory, not just the highest-visibility ones, and test whether documented controls are technically enforced for each
- 4Separate every finding into "not documented," "documented but not enforced," and "documented and enforced," since these require different remediation
- 5Score each control domain independently rather than producing a single blended program score
- 6Report findings to the audit committee or board with domain-level maturity scores and a remediation owner assigned to each gap
The Artifacts
- —AI governance maturity model (control domain × maturity level, with concrete criteria for each cell)
- —Program audit sampling plan (how systems are selected for testing, not just self-reported)
- —Finding classification template (not documented / documented not enforced / documented and enforced)
- —Domain-level maturity scorecard for board or audit committee reporting
The Output
A domain-by-domain maturity assessment of the AI governance program, with enforcement gaps distinguished from documentation gaps, and a named owner for every finding.
Auditing the program is not the same as auditing a system
Most AI audit activity in practice tests one system against one set of requirements: does this model meet the EU AI Act's documentation standard, does this hiring tool pass a bias assessment. That work matters, but it answers a narrower question than a board or regulator asking for assurance over the governance program as a whole. Program auditing asks whether the mechanisms that are supposed to catch problems across every AI system actually work, not whether one system happens to pass.
This distinction matters because a program can pass every individual system audit and still be structurally weak, if the systems selected for audit are not representative, if the audit only checks documentation rather than enforcement, or if newly deployed systems never enter the audit cycle at all. A program audit has to test the audit process itself, not just its outputs.
Build a maturity model with testable criteria per domain
A useful audit needs concrete, testable maturity criteria for each control domain, not general descriptions like "processes are informal" versus "processes are mature." For a domain like monitoring, level 1 might mean no automated drift detection exists, level 3 might mean automated alerts exist but response is manual and undocumented, and level 5 might mean automated detection, a documented response runbook, and a closed-loop record of every alert and its resolution. Vague maturity language produces vague audit findings that are hard to act on.
Score domains independently. A program can have excellent documentation practices and weak monitoring at the same time, and averaging those into one number obscures exactly the gap the audit exists to find. Present maturity as a profile across domains, not a single grade.
Test enforcement, not just documentation
The single most common audit failure mode is treating a written policy as evidence that a control exists. A human-oversight checkpoint that a bug can bypass, or an approval gate that exists procedurally but was never backed by a documented accuracy threshold, will both look compliant on paper. Real incidents confirm this pattern repeatedly: content moderation systems where a technical bypass let automated enforcement reach thousands of users, and inventory systems that reached production without the validation their approval process assumed had happened.
A rigorous audit samples actual system behavior, not just policy documents, and classifies every finding into whether a control is undocumented, documented but unenforced, or documented and enforced. The middle category is where the real risk concentrates, and it is invisible to an audit that only reviews paperwork.
Sample beyond the systems everyone already trusts
Audits naturally gravitate toward the highest-visibility AI systems, the ones that already have executive attention and a paper trail. The systems most likely to have gaps are the ones added to the inventory recently, adopted by a single team without broad visibility, or embedded inside a vendor product where the organization does not control the underlying model. A sampling plan that draws from the full inventory, weighted toward newer and lower-visibility systems, will surface gaps that a review of only the flagship deployments will miss entirely.
Related frameworks
Not sure where to start? Answer 3 questions and get a tailored compliance action plan.
What applies to me? →More guidance like this, every week
New playbook articles, governance controls, and the regulatory changes driving them. Every Thursday.
