AI Governance Institute
← News

Claude Opus 5.5 Brings Behavioral Changes That Require Governance Review

What happened

Anthropic published Prompting Claude Opus 5.5, a technical guide documenting behavioral differences between Claude Opus 5.5 and its predecessor. Extended thinking is on by default. The model reasons through problems before producing an answer. Teams must now configure effort levels and token limits explicitly to manage cost and latency. The documentation flags that Claude Opus 5.5 handles safety refusals differently. It is more cautious during unattended agentic runs where no human is present to intervene. It also treats pasted user text as a potential source of instructions that could override system prompts. These changes follow a pattern of Anthropic disclosing model-level behavioral shifts that carry direct compliance implications, as reported earlier in Claude Opus 5.5 governance implications. Organizations that deployed Claude Opus 5 under existing system prompt configurations cannot assume those configurations carry forward without review.

Why it matters

  • ·Unattended agentic runs now trigger more cautious model behavior by default. Teams that rely on Claude to complete tasks without human oversight must test whether existing workflows still complete as intended. The model may now pause or refuse steps that previously ran automatically.
  • ·Pasted user text is explicitly identified as a prompt injection risk, meaning content a user pastes into a conversation could instruct the model to ignore system-level rules. This is a direct threat to operator controls in customer-facing deployments and must be addressed through input validation before rollout.
  • ·Extended thinking is always on, and token limits must be set deliberately or costs can increase significantly. Compliance programs that approved Claude Opus 5 under a specific cost and output model must reassess those approvals. The operational profile of the new version is materially different.

Governance controls affected

What to do now

  • ☐Inventory every deployment where Claude Opus 5 is in use and flag it for re-review before migrating to Claude Opus 5.5.
  • ☐Ask your engineering team whether any Claude-powered workflows run without a human available to review or approve steps, and test whether those workflows still complete correctly under the new model's more cautious unattended behavior.
  • ☐Require your engineering team to set explicit maximum token limits for every Claude Opus 5.5 deployment to prevent unexpected cost increases from always-on extended thinking.
  • ☐Review all customer-facing Claude deployments where users can paste external text into the conversation, and confirm that input validation controls are in place to prevent that text from overriding system-level instructions.
  • ☐Update your AI model change documentation to reflect the behavioral differences between Claude Opus 5 and Claude Opus 5.5, and confirm that any prior compliance approvals are re-evaluated against the new model profile.

What to watch next

Anthropic has been documenting behavioral shifts with each major Claude release. Compliance teams should expect further guidance as Claude Opus 5.5 is deployed at scale and edge cases emerge. Teams operating in jurisdictions with formal model change management requirements should take note. This includes those preparing for obligations under the EU AI Act Implementation Timeline. Treat this release as a change event that may trigger reassessment under existing approval workflows. The 60-80% attack success rate against Claude Code Auto Mode reported earlier this year makes prompt injection defenses in the new model worth independent testing before production rollout.

Stay ahead of stories like this

Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.

Powered by Buttondown.

Related Coverage

Research2026-09-24

CSA Research: Indirect Prompt Injection Defeats AI Coding Agent Safety Classifier

A Cloud Security Alliance briefing published September 8, 2026 documents research showing indirect prompt injection defeating the safety classifier of an AI coding agent in a high proportion of controlled trials. The finding directly contradicts stronger vendor safety claims. Compliance teams governing agentic developer tools face an immediate gap between vendor assurances and independently verified runtime behavior.

Corporate Policy2026-09-23

Claude Opus 5.5 Cuts Cost 40% and Adds Dual-Use Verification Programs

Anthropic released Claude Opus 5.5 on September 22, 2026, a frontier model priced 40% below its predecessor. The release introduces formal verification programs for biology and cybersecurity use cases. External evaluators including Frontier Design and METR tested the model before release.

Research2026-09-19

BragJack Attack Turns Browser Extensions Into AI Agent Hijack Tools

Security researcher Gal Weizman disclosed a new attack class called BragJack, showing how a single malicious browser extension can seize control of AI agents in Chrome, Edge, Perplexity Comet, Opera Neon, and Claude for Chrome. Using a native browser mechanism, attackers can force hijacked agents to read local files, capture screenshots, access browsing history, and send emails on behalf of victims. Enterprise compliance programs are directly affected because the attacks exploit privileged AI agent access, not conventional malware, complicating detection and existing endpoint controls.