Multi-Agent Trust Hierarchy
Added May 2026
Define explicit rules for which AI agents can instruct, call on, or delegate authority to other agents in multi-agent systems.
Objective
Prevent privilege escalation (agents gaining powers they were never given) and unauthorized action chains in multi-agent systems by enforcing a documented, auditable trust model.
Maturity Levels
Initial
No trust hierarchy exists; agents can call on other agents without restriction.
Developing
Trust relationships are informally understood by the engineering team but not documented or enforced while the system runs.
Defined
A documented trust model specifies which agents can instruct which others, with automatic checks enforced while the system runs.
Managed
Decisions on which agents may direct others are logged and reviewed; unusual delegation patterns trigger alerts.
Optimizing
Trust relationships are checked automatically using tamper-proof digital proof of each agent's identity; trust model is reviewed after every change to the system design.
Evidence Requirements
What an auditor or assessor would expect to see for this control.
- —Documented trust model specifying which agents may call on which others and what actions each may accept from each caller
- —Logs of calls between agents with sender ID, recipient ID, instruction summary, timestamp, and resulting action for a sample period
- —Alert or investigation records for unusual patterns of agents delegating authority, detected in live use
- —System design review records confirming the trust model was reviewed and updated after each change to the system design
- —Configuration or test evidence confirming agents cannot grant their own permissions to subagents without human or system approval
Implementation Notes
Key steps
- Treat instructions passed between agents with the same skepticism as outside user inputs. An orchestrator agent (one that directs others) can be compromised or hijacked by prompt injection (hidden instructions planted in content it reads), so subagents should not blindly follow it.
- Implement an allowlist (an approved list): each agent explicitly lists which other agents may call on it and what actions it will accept from each.
- Log all calls between agents with the initiating agent's identity, the instruction passed, and the resulting action. This is essential for reconstructing events after an incident.
- Avoid agents that can grant their own permissions to subagents; privilege escalation must require human or system approval.
Example Implementation
Software engineering pipeline using an orchestrator and specialist sub-agents
Multi-Agent Trust Policy: Development Pipeline
| Agent | Role | May Invoke | May Accept Instructions From |
|---|---|---|---|
| Orchestrator | Task decomposition and assignment | Code Writer, Reviewer, Test Runner | Human operator only |
| Code Writer | Feature implementation | None | Orchestrator only |
| Code Reviewer | Diff review and feedback | None | Orchestrator only |
| Test Runner | Test execution and reporting | None | Orchestrator only |
Enforcement rules:
- No agent may accept task instructions from an agent at equal or lower trust level
- The Orchestrator may not grant its own permissions or those of sub-agents at runtime; any capability expansion requires human approval
- Any instruction that would cause an agent to act outside its permitted scope must cause a halt and route to human review
Audit requirement: All inter-agent invocations logged with sender_id, recipient_id, instruction summary, timestamp, and resulting action
