Kill-Switch Propagation Testing
Added May 2026
Test emergency stops across subagents and agents running in parallel. Verify that all agent activity stops within a defined time window.
Objective
Ensure the emergency halt capability from AGT-008 works in multi-agent deployments, where stop signals must pass along chains of agents handing off work and reach agents running simultaneously.
Maturity Levels
Initial
Kill-switch exists but propagation to subagents is untested; behavior when agents run in parallel is unknown.
Developing
Kill-switch propagation tested manually for simple single-agent deployments but not for multi-agent or parallel configurations.
Defined
Propagation tests are documented and executed for all production agent topologies on a scheduled basis; results are recorded.
Managed
Propagation latency is measured and tracked against an SLA; failures trigger remediation with a documented root cause.
Optimizing
Propagation tests run automatically on every topology change; limits on time to stop are enforced in the automated release pipeline (CI/CD) before agent deployment.
Evidence Requirements
What an auditor or assessor would expect to see for this control.
- —Agent topology map documenting all agent relationships, delegation depths, and parallel execution paths in production
- —Kill-switch propagation test plan with defined SLA (maximum latency from halt command to full cessation)
- —Executed test results showing propagation latency, pass/fail status, and coverage of all production topologies
- —Remediation records for any propagation failures, including root cause and fix verification
- —Change-trigger log showing propagation tests were re-executed after topology changes
Implementation Notes
Key steps
- Map how agents connect in production (the agent topology): document which agents can launch subagents, which run in parallel, and how many layers deep delegation can go.
- Define a propagation service level agreement (SLA): the maximum time from issuing a halt command to confirmed stopping of all agent activity across every agent and system.
- Write test cases that cover: (1) single-agent halt, (2) a lead agent passing the stop to its subagents, (3) parallel agent halt, (4) halt while an agent is using a tool or taking an irreversible action.
- Run tests on a schedule (at least quarterly, or after any topology change) and record pass/fail, measured latency (time to stop), and any agents that failed to halt.
- For agents that interact with external systems, verify that requests already sent to those systems are either completed cleanly or reversed, not left half-finished.
Example Implementation
Enterprise deploying a customer-service orchestrator that spawns research, drafting, and CRM-update subagents
Kill-Switch Propagation Test Results: Q2 2026
Topology tested: Orchestrator → [Research Agent, Drafting Agent] → CRM Update Agent
| Test case | Halt command issued | All agents stopped | Latency | In-flight actions |
|---|---|---|---|---|
| Single agent (Drafting) | 14:02:01 | 14:02:01 | 0.3s | Completed cleanly |
| Orchestrator → 2 subagents | 14:05:00 | 14:05:02 | 2.1s | Research cancelled, Drafting completed |
| Halt during CRM write | 14:08:30 | 14:08:32 | 1.8s | CRM write rolled back |
| Parallel execution (3 agents) | 14:11:00 | 14:11:03 | 2.7s | All stopped |
SLA: 5 seconds. All tests passed. Finding: CRM rollback requires manual verification, add automated confirmation check.
