AI Governance Institute
All governance templates →How do we ensure human-in-the-loop review is actually effective?

Implementation Kit

AI Human Review Audit Template and Override Analysis

Auditing a review process you already have for whether it is substantive or a formality. A workflow audit template, an override-rate audit by reviewer, and second-reviewer escalation criteria for cases where the AI and the first reviewer diverge.

Who this is for: The risk or audit function checking whether existing human-in-the-loop review actually reduces risk.

Download the kit (Markdown) ↓3 artifacts. Every table also copies as CSV.

1. Human review workflow audit

Spreadsheet

One row per AI system with a human step. Records what the reviewer actually has.

Template

SystemInfo available to reviewerTime allocated per caseOverride mechanismReasoning documented?Auditor verdict
<system>AI output only / output + factors / full recordseconds/minutesone-click / escalation required / noneevery case / overrides only / noneSubstantive / Formality / Mixed

Worked example

SystemInfo availableTime per caseOverride mechanismReasoning documented?Verdict
resume-screeneroutput + top factors45s avgone-clickoverrides onlyFormality: too fast, no agree-rationale, auto-advance
fraud-holdfull transaction record + model reasons3 min avgone-clickevery caseSubstantive
refund-copilotdraft text only20sedit-then-sendnoneFormality: reviewer edits wording, not the decision

Acceptance criteria

  • Every AI system with a human review step is audited, not a sample.
  • Time-per-case comes from logs, not from what the process owner says it should be.
  • Each system gets an explicit verdict, and "Formality" or "Mixed" opens a remediation item.

2. Override-rate audit by reviewer

Spreadsheet

Finds rubber-stamping. Near-zero override rates and near-instant reviews are the signals.

Template

ReviewerReviews (period)Override rateMedian time10th-pct timeAppeal reversal rateFlag
<id>%%none / low-override / too-fast / high-reversal

Worked example

ReviewerReviewsOverride rateMedian time10th-pct timeAppeal reversalFlag
u-33902,1400.6%14s6s3.9%low-override + too-fast + high-reversal
u-44711,88012%3m 10s1m 40s1.8%none
team median11%2m 50s2.1%

Acceptance criteria

  • Rates are computed per reviewer against a team median or a calibrated range.
  • A flagged reviewer gets a workload and tooling review before any performance conversation.
  • Appeal reversal rate is included, so speed is weighed against accuracy.

3. Second-reviewer escalation criteria

Spreadsheet

When a case needs a second person: AI and first reviewer disagree, or stakes are high.

Template

TriggerThresholdSecond reviewerSLA
AI and first reviewer diverge sharplyAI score in top/bottom band and reviewer flips itsenior reviewer, different person<time>
High-stakes outcomedenial, termination, account closuresupervisor sign-off<time>
Novel or edge casereviewer marks "uncertain"domain specialist<time>
Pattern flagreviewer's recent decisions flagged by the override auditQA reviewer, sample<cadence>

Worked example

TriggerThresholdSecond reviewerSLA
Sharp divergenceAI says "advance", reviewer rejects a top-quartile score (or vice versa)senior recruitersame day
High-stakesany auto-decline pathnot applicable: no auto-decline; all declines are human + supervisor2 days
Edge case"uncertain" flagtalent partner2 days
Pattern flagreviewer flagged by override auditQA samples 10% for 4 weeksweekly

Acceptance criteria

  • The divergence trigger is defined concretely, not "significant disagreement".
  • High-stakes outcomes always get a second person.
  • Second-review outcomes feed back into reviewer coaching and the workflow audit.

Governance controls this kit produces evidence for

Completing the artifacts above gives you a head start on the evidence requirements for these controls.

HOC-003
HOC-003

The workflow audit is an assessment of the output review workflow's effectiveness.

HOC-004
HOC-004

The override-rate audit is the automation-bias monitoring evidence.

HOC-005
HOC-005

Flagged reviewers and coaching tie back to reviewer competency requirements.

HOC-006
HOC-006

The second-reviewer criteria are the escalation procedure for contested cases.

MON-004
MON-004

Appeal-reversal and divergence tracking are output-anomaly signals for the review process.

This kit backs one playbook. Read the full guidance for the reasoning behind each artifact.

Decide what to implement next

Assess your governance gaps, then create an action plan with owners and target dates. Build and export without an account; sign in when you want to save your plan.

Start the AI governance assessment →