Human Approval Gate for Irreversible Agent Actions
Added May 2026
Require explicit human approval before hard-to-reverse AI agent actions. Examples include communications, record changes, transactions, and data deletion.
Objective
Ensure humans retain control over consequential agent actions, preventing costly or harmful mistakes that cannot be undone automatically.
Maturity Levels
Initial
Agents execute all actions autonomously without approval gates.
Developing
Some high-risk actions require approval but the list is incomplete and not systematically maintained.
Defined
A documented list of action types requiring approval is enforced by the software that runs the agent, not just by convention.
Managed
Approval gate effectiveness is monitored; the time taken to approve is tracked; queued actions that expire without approval are flagged.
Optimizing
The list of gated actions is continuously refined based on incident data; low-risk approved actions are progressively delegated back to the agent.
Evidence Requirements
What an auditor or assessor would expect to see for this control.
- —Versioned, approved list of gated action types with rationale for inclusion
- —Evidence from the agent framework (code review, configuration) confirming gates are enforced before any tool runs, not solely by model instruction
- —Approval queue records showing pending, approved, and expired approval requests with timestamps
- —Records of actions cancelled after timeout expiry and how they were re-queued or closed
- —Reviewer screen specification or screenshot evidence confirming full action context is presented before approval
Implementation Notes
Key steps
- Define irreversible action categories before deployment: sending messages, creating/modifying/deleting records, financial transactions, code deployments, and requests to outside systems (API calls) that trigger real-world effects.
- Build approval gates into the agent framework (the software that carries out the agent's actions) so they apply before any tool runs, not as instructions to the model. Model instructions can be overridden by prompt injection (hidden commands planted in content the agent reads) or by reasoning errors.
- Design approval screens to show the full action context, not just the action itself, so reviewers can assess whether it is appropriate.
- Set a timeout for pending approvals, because an old approval is dangerous once the situation has changed; require re-confirmation after a defined window.
Example Implementation
Marketing team using an AI agent to draft and send campaign emails and update CRM records
Irreversible Action Gate Configuration: Campaign Agent
Gated action types (human approval required before execution):
- Send email to external recipients
- Mass-update CRM contact records (> 10 records)
- Create or modify campaign workflows
- Publish content to external channels
- Delete any records
Non-gated actions (agent may execute autonomously):
- Read CRM data
- Draft email copy (stored as draft, not sent)
- Generate audience segment previews
- Create internal notes
Approval timeout: 2 hours, if approval is not confirmed within 2 hours, action is cancelled and re-queued for the next business day
Approval UI requirement: Reviewer must see the full action context (recipient list preview, email content, affected records) before confirming; summary-only views are not permitted
