Red-Team Results Don't Transfer Across Agent Harnesses, NHIMG Finds
What happened
The NHI Management Group published Agent harness choice materially changes autonomous red teaming outcomes, an analysis finding that the middleware, gateway, and tool connections wrapped around an autonomous agent are co-determinants of its safety behavior, not neutral scaffolding. When the same underlying model is evaluated through different harnesses, the results diverge in ways that make harness-specific test outcomes non-transferable to other deployment configurations. The analysis connects directly to an earlier NHIMG finding covered in NHIMG: Agentic AI Governance Must Shift to Action-Level Runtime Controls, which similarly argued that model-level controls are insufficient for agentic deployments. The paper recommends three concrete governance changes: standardizing a set of approved harnesses for production use, restricting tool exposure to permissions scoped strictly to the task at hand, and benchmarking the model and harness together as a unified deployment unit for pre-deployment approval and ongoing monitoring purposes. The implication for compliance programs is significant: any red-team or safety evaluation conducted on a model without fixing the full harness configuration may not be a valid basis for a deployment approval decision.
Why it matters
- ·Pre-deployment approval gates that evaluate only the underlying model, without specifying and locking the harness, middleware, and tool set, may produce safety findings that do not hold in production, exposing organizations to unvalidated risk at the moment of go-live.
- ·Tool-scoping and permission boundary controls become meaningless if they are not assessed as part of the same evaluation unit as the model; organizations that have approved agentic deployments based on model-only red-teaming should treat those approvals as potentially requiring reassessment.
- ·Audit and incident response programs that log model behavior without capturing harness configuration at the time of an incident will lack the forensic detail needed to reconstruct whether a harmful output was a model failure or a harness configuration failure, complicating post-incident review and any regulatory disclosure.
Governance controls affected
What to do now
- ☐Audit all existing agentic deployment approvals to determine whether red-team evaluations were conducted against a fixed, production-equivalent harness configuration or against the model in isolation.
- ☐Update the pre-deployment approval gate policy to require that the model and its full harness stack, including middleware, gateway, and tool connections, be evaluated together as a single unit before any agentic deployment is approved.
- ☐Establish a registry of approved harnesses and require that any change to harness configuration, not just model version changes, triggers a new pre-deployment review.
- ☐Revise tool-exposure policies to enforce task-scoped permissions at the harness level and verify that these restrictions are tested as part of each red-team exercise.
- ☐Update incident response and audit log standards to capture harness configuration state alongside model version at the time of any logged output or incident, so forensic reconstruction is possible.
What to watch next
Organizations should monitor whether emerging agentic governance standards, including the output of the ITU Focus Group on Trust and Identity for Humans and Agentic AI and future NHIMG publications, incorporate harness-level evaluation requirements into formal frameworks. The finding also raises questions for regulators considering pre-deployment testing mandates: if testing requirements do not specify harness configuration as part of the evaluated unit, those mandates may provide less assurance than intended. Further research on how harness variation interacts with tool-call permission drift, a risk highlighted in the Frontier Agents Fail Policy Tests at Scale coverage, would materially sharpen the governance picture for enterprise teams.
Stay ahead of stories like this
Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.
