Implementation Kit
AI Human Review Audit Template and Override Analysis
Auditing a review process you already have for whether it is substantive or a formality. A workflow audit template, an override-rate audit by reviewer, and second-reviewer escalation criteria for cases where the AI and the first reviewer diverge.
Who this is for: The risk or audit function checking whether existing human-in-the-loop review actually reduces risk.
1. Human review workflow audit
SpreadsheetOne row per AI system with a human step. Records what the reviewer actually has.
Template
| System | Info available to reviewer | Time allocated per case | Override mechanism | Reasoning documented? | Auditor verdict |
|---|---|---|---|---|---|
| <system> | AI output only / output + factors / full record | seconds/minutes | one-click / escalation required / none | every case / overrides only / none | Substantive / Formality / Mixed |
Worked example
| System | Info available | Time per case | Override mechanism | Reasoning documented? | Verdict |
|---|---|---|---|---|---|
| resume-screener | output + top factors | 45s avg | one-click | overrides only | Formality: too fast, no agree-rationale, auto-advance |
| fraud-hold | full transaction record + model reasons | 3 min avg | one-click | every case | Substantive |
| refund-copilot | draft text only | 20s | edit-then-send | none | Formality: reviewer edits wording, not the decision |
Acceptance criteria
- ✓Every AI system with a human review step is audited, not a sample.
- ✓Time-per-case comes from logs, not from what the process owner says it should be.
- ✓Each system gets an explicit verdict, and "Formality" or "Mixed" opens a remediation item.
2. Override-rate audit by reviewer
SpreadsheetFinds rubber-stamping. Near-zero override rates and near-instant reviews are the signals.
Template
| Reviewer | Reviews (period) | Override rate | Median time | 10th-pct time | Appeal reversal rate | Flag |
|---|---|---|---|---|---|---|
| <id> | % | % | none / low-override / too-fast / high-reversal |
Worked example
| Reviewer | Reviews | Override rate | Median time | 10th-pct time | Appeal reversal | Flag |
|---|---|---|---|---|---|---|
| u-3390 | 2,140 | 0.6% | 14s | 6s | 3.9% | low-override + too-fast + high-reversal |
| u-4471 | 1,880 | 12% | 3m 10s | 1m 40s | 1.8% | none |
| team median | 11% | 2m 50s | 2.1% |
Acceptance criteria
- ✓Rates are computed per reviewer against a team median or a calibrated range.
- ✓A flagged reviewer gets a workload and tooling review before any performance conversation.
- ✓Appeal reversal rate is included, so speed is weighed against accuracy.
3. Second-reviewer escalation criteria
SpreadsheetWhen a case needs a second person: AI and first reviewer disagree, or stakes are high.
Template
| Trigger | Threshold | Second reviewer | SLA |
|---|---|---|---|
| AI and first reviewer diverge sharply | AI score in top/bottom band and reviewer flips it | senior reviewer, different person | <time> |
| High-stakes outcome | denial, termination, account closure | supervisor sign-off | <time> |
| Novel or edge case | reviewer marks "uncertain" | domain specialist | <time> |
| Pattern flag | reviewer's recent decisions flagged by the override audit | QA reviewer, sample | <cadence> |
Worked example
| Trigger | Threshold | Second reviewer | SLA |
|---|---|---|---|
| Sharp divergence | AI says "advance", reviewer rejects a top-quartile score (or vice versa) | senior recruiter | same day |
| High-stakes | any auto-decline path | not applicable: no auto-decline; all declines are human + supervisor | 2 days |
| Edge case | "uncertain" flag | talent partner | 2 days |
| Pattern flag | reviewer flagged by override audit | QA samples 10% for 4 weeks | weekly |
Acceptance criteria
- ✓The divergence trigger is defined concretely, not "significant disagreement".
- ✓High-stakes outcomes always get a second person.
- ✓Second-review outcomes feed back into reviewer coaching and the workflow audit.
Governance controls this kit produces evidence for
Completing the artifacts above gives you a head start on the evidence requirements for these controls.
The workflow audit is an assessment of the output review workflow's effectiveness.
The override-rate audit is the automation-bias monitoring evidence.
Flagged reviewers and coaching tie back to reviewer competency requirements.
The second-reviewer criteria are the escalation procedure for contested cases.
Appeal-reversal and divergence tracking are output-anomaly signals for the review process.
This kit backs one playbook. Read the full guidance for the reasoning behind each artifact.
Decide what to implement next
Assess your governance gaps, then create an action plan with owners and target dates. Build and export without an account; sign in when you want to save your plan.
Start the AI governance assessment →