# AI Human Review Audit Template and Override Analysis

Auditing a review process you already have for whether it is substantive or a formality. A workflow audit template, an override-rate audit by reviewer, and second-reviewer escalation criteria for cases where the AI and the first reviewer diverge.

**Who this is for:** The risk or audit function checking whether existing human-in-the-loop review actually reduces risk.

Source playbook: https://aigovernance.com/playbook/effective-human-in-the-loop

---

## Human review workflow audit

_One row per AI system with a human step. Records what the reviewer actually has._

### Template

| System | Info available to reviewer | Time allocated per case | Override mechanism | Reasoning documented? | Auditor verdict |
|---|---|---|---|---|---|
| <system> | AI output only / output + factors / full record | seconds/minutes | one-click / escalation required / none | every case / overrides only / none | Substantive / Formality / Mixed |

### Worked example

| System | Info available | Time per case | Override mechanism | Reasoning documented? | Verdict |
|---|---|---|---|---|---|
| resume-screener | output + top factors | 45s avg | one-click | overrides only | Formality: too fast, no agree-rationale, auto-advance |
| fraud-hold | full transaction record + model reasons | 3 min avg | one-click | every case | Substantive |
| refund-copilot | draft text only | 20s | edit-then-send | none | Formality: reviewer edits wording, not the decision |

### Acceptance criteria

- Every AI system with a human review step is audited, not a sample.
- Time-per-case comes from logs, not from what the process owner says it should be.
- Each system gets an explicit verdict, and "Formality" or "Mixed" opens a remediation item.

---

## Override-rate audit by reviewer

_Finds rubber-stamping. Near-zero override rates and near-instant reviews are the signals._

### Template

| Reviewer | Reviews (period) | Override rate | Median time | 10th-pct time | Appeal reversal rate | Flag |
|---|---|---|---|---|---|---|
| <id> | | % | | | % | none / low-override / too-fast / high-reversal |

### Worked example

| Reviewer | Reviews | Override rate | Median time | 10th-pct time | Appeal reversal | Flag |
|---|---|---|---|---|---|---|
| u-3390 | 2,140 | 0.6% | 14s | 6s | 3.9% | low-override + too-fast + high-reversal |
| u-4471 | 1,880 | 12% | 3m 10s | 1m 40s | 1.8% | none |
| team median | | 11% | 2m 50s | | 2.1% | |

### Acceptance criteria

- Rates are computed per reviewer against a team median or a calibrated range.
- A flagged reviewer gets a workload and tooling review before any performance conversation.
- Appeal reversal rate is included, so speed is weighed against accuracy.

---

## Second-reviewer escalation criteria

_When a case needs a second person: AI and first reviewer disagree, or stakes are high._

### Template

| Trigger | Threshold | Second reviewer | SLA |
|---|---|---|---|
| AI and first reviewer diverge sharply | AI score in top/bottom band and reviewer flips it | senior reviewer, different person | <time> |
| High-stakes outcome | denial, termination, account closure | supervisor sign-off | <time> |
| Novel or edge case | reviewer marks "uncertain" | domain specialist | <time> |
| Pattern flag | reviewer's recent decisions flagged by the override audit | QA reviewer, sample | <cadence> |

### Worked example

| Trigger | Threshold | Second reviewer | SLA |
|---|---|---|---|
| Sharp divergence | AI says "advance", reviewer rejects a top-quartile score (or vice versa) | senior recruiter | same day |
| High-stakes | any auto-decline path | not applicable: no auto-decline; all declines are human + supervisor | 2 days |
| Edge case | "uncertain" flag | talent partner | 2 days |
| Pattern flag | reviewer flagged by override audit | QA samples 10% for 4 weeks | weekly |

### Acceptance criteria

- The divergence trigger is defined concretely, not "significant disagreement".
- High-stakes outcomes always get a second person.
- Second-review outcomes feed back into reviewer coaching and the workflow audit.

---

## Governance controls this kit produces evidence for

- **HOC-003**: The workflow audit is an assessment of the output review workflow's effectiveness.
- **HOC-004**: The override-rate audit is the automation-bias monitoring evidence.
- **HOC-005**: Flagged reviewers and coaching tie back to reviewer competency requirements.
- **HOC-006**: The second-reviewer criteria are the escalation procedure for contested cases.
- **MON-004**: Appeal-reversal and divergence tracking are output-anomaly signals for the review process.
