AI Governance Institute
All governance templates →How do we audit our AI governance program, not just individual AI systems?

Implementation Kit

AI Governance Program Audit Templates and Maturity Scorecard

Auditing the governance program itself, not just individual systems. A maturity model with concrete criteria per cell, a sampling plan that does not rely on self-report, a three-state finding classification, and a domain scorecard for the board.

Who this is for: Internal audit or an external assessor giving an opinion on whether the AI governance program works.

Download the kit (Markdown) ↓4 artifacts. Every table also copies as CSV.

1. AI governance maturity model

Spreadsheet

Control domain by maturity level, with a concrete, testable criterion in each cell.

Template

Domain1 Ad hoc2 Defined3 Enforced4 Optimized
Inventory and classificationsome systems listeddocumented method; most systems tieredall systems tiered; intake gate enforcedintake automated; drift from the register alerted
Risk assessmentdone for some systemstemplate exists; done pre-deployment for high-risksigned assessments for all high-risk; gaps block deploymentassessments partly automated; linked to monitoring
Human oversightreviewers existdesign documentedoverride monitoring; reviewers qualifiedoversight effectiveness measured and improved
Monitoringmanual spot checksmetrics definedthresholds + alerting in productionauto-response for defined conditions
Incident responseinformalrunbook + registerdrills run; notifications assessed every timelessons systematically fed back into controls
Third-partycontracts reviewed ad hocDD process definedall material vendors assessed + monitoredrequalification automated on model change
Regulatorysomeone watches the newssources + owners definedobligations mapped to controls; calendar maintainedhorizon scanning drives roadmap
Board reportingoccasionaltemplate + cadencereconciled to source; decisions requestedboard literacy measured; oversight effectiveness reviewed

Worked example

Assessed against the program, H1 2026:

DomainScoreEvidence
Inventory and classification2 (heading to 3)method documented; 2 systems found off-register in audit
Risk assessment2template exists; 1 of 3 high-risk unsigned
Human oversight1design docs scattered; effectiveness gap confirmed by interviews
Monitoring1metrics defined for 2 systems; no alerting
Incident response2runbook + register; one drill run
Third-party2DD process defined; monitoring partial
Regulatory2sources + owners; mapping in progress
Board reporting2template + cadence; not yet reconciled to source

Acceptance criteria

  • Every cell has a concrete, testable criterion, not an adjective.
  • Scores are assigned from audit evidence, not management's self-rating.
  • The model is stable across audits so trend is meaningful.

2. Program audit sampling plan

Spreadsheet

How systems and controls are selected for testing, so coverage is not owner-curated.

Template

PopulationSelection methodSample sizeRationale
AI systems for control testingrandom from the full register, stratified by tier (all High, sample of Limited/Minimal)tier drives obligation
Decisions for reconstructionrandom by seed from the full decision log of sampled systemstests the audit trail
Reviewers for interviewrandom from the active reviewer pool, not manager-nominatedtests oversight in practice
Vendor DD recordsall material vendors + random sample of the restmateriality plus coverage
Incidents for tracingall Sev-1/Sev-2 + sample of Sev-3 and near-misseslearn from the small ones too

Worked example

H1 plan as executed.

PopulationSample drawn
AI systems for control testingall 3 High + 5 of 44 others (random); 8 controls tested
Decisions for reconstruction25 per sampled High system, auditor-seeded
Reviewers for interview5 of 22, random
Vendor DD records4 material + 3 random
Incidents for tracingall 2 Sev-2 + 3 of 9 near-misses

Independent discovery for inventory completeness run in parallel.

Acceptance criteria

  • Samples are drawn by the auditor, using randomisation, not provided by the process owner.
  • High-tier systems and serious incidents are fully covered, not sampled.
  • Inventory completeness is tested by independent discovery, not by reviewing the register.

3. Finding classification template

Spreadsheet

Three states that separate "we wrote it down" from "it actually works".

Template

StateMeaningExampleTypical action
Not documentedno policy, standard, or defined process existsno monitoring standardwrite it; assign an owner
Documented, not enforcedit exists on paper but system behaviour or records show it is not consistently applieddrift monitoring standard exists; 2 of 3 models have no alertingclose the gap between paper and practice; add a check
Documented and enforcedexists and testing confirms it operatesdecision logging standard; sampled decisions all fully loggedmaintain; consider optimisation

Worked example

H1 findings grouped by state.

StateFindings
Not documentedmonitoring alerting standard; second-reviewer protocol
Documented, not enforcedhuman oversight (design exists, practice is a formality); onboarding intake gate (2 systems bypassed it)
Documented and enforceddecision logging; bias testing protocol; vendor DD for material vendors

Acceptance criteria

  • Every finding is placed in one of the three states with the evidence for that state.
  • "Documented, not enforced" findings cite the behaviour or records that show the gap.
  • The split between states is reported to the board, not just a finding count.

4. Domain-level maturity scorecard

Spreadsheet

The one-page view for the board or audit committee: score per domain, trend, and target.

Template

DomainScore (0-4)Prior auditTargetTrendKey gap
<domain>up / flat / down

Worked example

DomainScorePriorTargetTrendKey gap
Inventory and classification213upintake gate bypassed twice
Risk assessment213upone High-tier system unsigned
Human oversight112flatoversight is a formality in practice
Monitoring102upno alerting
Incident response212upmet target
Third-party223flatmonitoring cadence not enforced
Regulatory213upcontrol mapping incomplete
Board reporting212upnot reconciled to source
Overall1.751.02.5upoversight and monitoring

Acceptance criteria

  • The scorecard shows score, prior, target, and trend per domain on one page.
  • It names the single key gap per domain.
  • Overall score reconciles with the per-domain scores.

Governance controls this kit produces evidence for

Completing the artifacts above gives you a head start on the evidence requirements for these controls.

MGV-004
MGV-004

The whole kit is the continuous AI assurance function's program-level audit method.

BRD-005
BRD-005

The maturity model and scorecard are the governance maturity assessment.

MGV-003
MGV-003

Findings and targets become governance-program milestones.

HOC-007
HOC-007

The domain scorecard is board and audit-committee reporting on program maturity.

BRD-002
BRD-002

Audit independence and reporting lines reflect the committee charter's assurance expectations.

This kit backs one playbook. Read the full guidance for the reasoning behind each artifact.

Decide what to implement next

Assess your governance gaps, then create an action plan with owners and target dates. Build and export without an account; sign in when you want to save your plan.

Start the AI governance assessment →