Intelligence for Compliance and GRC Teams
1 item
Researchers have published HANDBOOK.md, a benchmark of 65 agentic tasks testing whether language model agents follow long-form enterprise policy documents during extended tool use. Under strict grading, the best-performing model configuration passed only 36.2% of trials, with most frontier models falling below 25%. Failure patterns include agents overriding standing policy in response to in-context requests, acting against completed compliance checks, and losing rule details over long task horizons.