# AI Governance Program Audit Templates and Maturity Scorecard

Auditing the governance program itself, not just individual systems. A maturity model with concrete criteria per cell, a sampling plan that does not rely on self-report, a three-state finding classification, and a domain scorecard for the board.

**Who this is for:** Internal audit or an external assessor giving an opinion on whether the AI governance program works.

Source playbook: https://aigovernance.com/playbook/ai-governance-auditing

---

## AI governance maturity model

_Control domain by maturity level, with a concrete, testable criterion in each cell._

### Template

| Domain | 1 Ad hoc | 2 Defined | 3 Enforced | 4 Optimized |
|---|---|---|---|---|
| Inventory and classification | some systems listed | documented method; most systems tiered | all systems tiered; intake gate enforced | intake automated; drift from the register alerted |
| Risk assessment | done for some systems | template exists; done pre-deployment for high-risk | signed assessments for all high-risk; gaps block deployment | assessments partly automated; linked to monitoring |
| Human oversight | reviewers exist | design documented | override monitoring; reviewers qualified | oversight effectiveness measured and improved |
| Monitoring | manual spot checks | metrics defined | thresholds + alerting in production | auto-response for defined conditions |
| Incident response | informal | runbook + register | drills run; notifications assessed every time | lessons systematically fed back into controls |
| Third-party | contracts reviewed ad hoc | DD process defined | all material vendors assessed + monitored | requalification automated on model change |
| Regulatory | someone watches the news | sources + owners defined | obligations mapped to controls; calendar maintained | horizon scanning drives roadmap |
| Board reporting | occasional | template + cadence | reconciled to source; decisions requested | board literacy measured; oversight effectiveness reviewed |

### Worked example

**Assessed against the program, H1 2026:**
| Domain | Score | Evidence |
|---|---|---|
| Inventory and classification | 2 (heading to 3) | method documented; 2 systems found off-register in audit |
| Risk assessment | 2 | template exists; 1 of 3 high-risk unsigned |
| Human oversight | 1 | design docs scattered; effectiveness gap confirmed by interviews |
| Monitoring | 1 | metrics defined for 2 systems; no alerting |
| Incident response | 2 | runbook + register; one drill run |
| Third-party | 2 | DD process defined; monitoring partial |
| Regulatory | 2 | sources + owners; mapping in progress |
| Board reporting | 2 | template + cadence; not yet reconciled to source |

### Acceptance criteria

- Every cell has a concrete, testable criterion, not an adjective.
- Scores are assigned from audit evidence, not management's self-rating.
- The model is stable across audits so trend is meaningful.

---

## Program audit sampling plan

_How systems and controls are selected for testing, so coverage is not owner-curated._

### Template

| Population | Selection method | Sample size | Rationale |
|---|---|---|---|
| AI systems for control testing | random from the full register, stratified by tier (all High, sample of Limited/Minimal) | | tier drives obligation |
| Decisions for reconstruction | random by seed from the full decision log of sampled systems | | tests the audit trail |
| Reviewers for interview | random from the active reviewer pool, not manager-nominated | | tests oversight in practice |
| Vendor DD records | all material vendors + random sample of the rest | | materiality plus coverage |
| Incidents for tracing | all Sev-1/Sev-2 + sample of Sev-3 and near-misses | | learn from the small ones too |

### Worked example

> H1 plan as executed.

| Population | Sample drawn |
|---|---|
| AI systems for control testing | all 3 High + 5 of 44 others (random); 8 controls tested |
| Decisions for reconstruction | 25 per sampled High system, auditor-seeded |
| Reviewers for interview | 5 of 22, random |
| Vendor DD records | 4 material + 3 random |
| Incidents for tracing | all 2 Sev-2 + 3 of 9 near-misses |

Independent discovery for inventory completeness run in parallel.

### Acceptance criteria

- Samples are drawn by the auditor, using randomisation, not provided by the process owner.
- High-tier systems and serious incidents are fully covered, not sampled.
- Inventory completeness is tested by independent discovery, not by reviewing the register.

---

## Finding classification template

_Three states that separate "we wrote it down" from "it actually works"._

### Template

| State | Meaning | Example | Typical action |
|---|---|---|---|
| Not documented | no policy, standard, or defined process exists | no monitoring standard | write it; assign an owner |
| Documented, not enforced | it exists on paper but system behaviour or records show it is not consistently applied | drift monitoring standard exists; 2 of 3 models have no alerting | close the gap between paper and practice; add a check |
| Documented and enforced | exists and testing confirms it operates | decision logging standard; sampled decisions all fully logged | maintain; consider optimisation |

### Worked example

> H1 findings grouped by state.

| State | Findings |
|---|---|
| Not documented | monitoring alerting standard; second-reviewer protocol |
| Documented, not enforced | human oversight (design exists, practice is a formality); onboarding intake gate (2 systems bypassed it) |
| Documented and enforced | decision logging; bias testing protocol; vendor DD for material vendors |

### Acceptance criteria

- Every finding is placed in one of the three states with the evidence for that state.
- "Documented, not enforced" findings cite the behaviour or records that show the gap.
- The split between states is reported to the board, not just a finding count.

---

## Domain-level maturity scorecard

_The one-page view for the board or audit committee: score per domain, trend, and target._

### Template

| Domain | Score (0-4) | Prior audit | Target | Trend | Key gap |
|---|---|---|---|---|---|
| <domain> | | | | up / flat / down | |

### Worked example

| Domain | Score | Prior | Target | Trend | Key gap |
|---|---|---|---|---|---|
| Inventory and classification | 2 | 1 | 3 | up | intake gate bypassed twice |
| Risk assessment | 2 | 1 | 3 | up | one High-tier system unsigned |
| Human oversight | 1 | 1 | 2 | flat | oversight is a formality in practice |
| Monitoring | 1 | 0 | 2 | up | no alerting |
| Incident response | 2 | 1 | 2 | up | met target |
| Third-party | 2 | 2 | 3 | flat | monitoring cadence not enforced |
| Regulatory | 2 | 1 | 3 | up | control mapping incomplete |
| Board reporting | 2 | 1 | 2 | up | not reconciled to source |
| **Overall** | **1.75** | **1.0** | **2.5** | up | oversight and monitoring |

### Acceptance criteria

- The scorecard shows score, prior, target, and trend per domain on one page.
- It names the single key gap per domain.
- Overall score reconciles with the per-domain scores.

---

## Governance controls this kit produces evidence for

- **MGV-004**: The whole kit is the continuous AI assurance function's program-level audit method.
- **BRD-005**: The maturity model and scorecard are the governance maturity assessment.
- **MGV-003**: Findings and targets become governance-program milestones.
- **HOC-007**: The domain scorecard is board and audit-committee reporting on program maturity.
- **BRD-002**: Audit independence and reporting lines reflect the committee charter's assurance expectations.
