AI Governance Institute
← Agentic AI
AGT · Agentic AIAGT-005Medium effortAgent-relevant

Human Approval Gate for Irreversible Agent Actions

Added May 2026

Require explicit human approval before hard-to-reverse AI agent actions. Examples include communications, record changes, transactions, and data deletion.

Objective

Ensure humans retain control over consequential agent actions, preventing costly or harmful mistakes that cannot be undone automatically.

Maturity Levels

1

Initial

Agents execute all actions autonomously without approval gates.

2

Developing

Some high-risk actions require approval but the list is incomplete and not systematically maintained.

3

Defined

A documented list of action types requiring approval is enforced by the software that runs the agent, not just by convention.

4

Managed

Approval gate effectiveness is monitored; the time taken to approve is tracked; queued actions that expire without approval are flagged.

5

Optimizing

The list of gated actions is continuously refined based on incident data; low-risk approved actions are progressively delegated back to the agent.

Evidence Requirements

What an auditor or assessor would expect to see for this control.

  • —Versioned, approved list of gated action types with rationale for inclusion
  • —Evidence from the agent framework (code review, configuration) confirming gates are enforced before any tool runs, not solely by model instruction
  • —Approval queue records showing pending, approved, and expired approval requests with timestamps
  • —Records of actions cancelled after timeout expiry and how they were re-queued or closed
  • —Reviewer screen specification or screenshot evidence confirming full action context is presented before approval

Implementation Notes

Key steps

  • Define irreversible action categories before deployment: sending messages, creating/modifying/deleting records, financial transactions, code deployments, and requests to outside systems (API calls) that trigger real-world effects.
  • Build approval gates into the agent framework (the software that carries out the agent's actions) so they apply before any tool runs, not as instructions to the model. Model instructions can be overridden by prompt injection (hidden commands planted in content the agent reads) or by reasoning errors.
  • Design approval screens to show the full action context, not just the action itself, so reviewers can assess whether it is appropriate.
  • Set a timeout for pending approvals, because an old approval is dangerous once the situation has changed; require re-confirmation after a defined window.

Example Implementation

Marketing team using an AI agent to draft and send campaign emails and update CRM records

Irreversible Action Gate Configuration: Campaign Agent

Gated action types (human approval required before execution):

  • Send email to external recipients
  • Mass-update CRM contact records (> 10 records)
  • Create or modify campaign workflows
  • Publish content to external channels
  • Delete any records

Non-gated actions (agent may execute autonomously):

  • Read CRM data
  • Draft email copy (stored as draft, not sent)
  • Generate audience segment previews
  • Create internal notes

Approval timeout: 2 hours, if approval is not confirmed within 2 hours, action is cancelled and re-queued for the next business day

Approval UI requirement: Reviewer must see the full action context (recipient list preview, email content, affected records) before confirming; summary-only views are not permitted

Control Details

Control ID
AGT-005
Typical owner
AI Governance Team / AI Engineering
Implementation effort
Medium effort
Agent-relevant
Yes

Tags

human-in-the-loopirreversible actionsagent approvalagentic AI

Templates for this control

Get control updates weekly

New and updated controls, maturity guidance, and the regulatory changes behind them. Every Thursday.

Powered by Buttondown.