AI Governance Institute
← News

Safety Pause Commitment Dropped from Anthropic's Responsible Scaling Policy Version 3.0

What happened

Anthropic released version 3.0 of its Responsible Scaling Policy (RSP) in February 2026, removing the company's original commitment to pause AI development if safety could not be guaranteed in advance. The safety pause provision had been a defining feature of Anthropic's voluntary governance framework since the RSP was first introduced in 2023. The removal represents a material shift away from a precautionary halt mechanism, though the specific replacement controls introduced in version 3.0 have not been fully detailed in public reporting. The source reporting was published by AI Coachella Valley at a briefs index page rather than a dedicated article. Organizations that have cited Anthropic's published governance commitments in vendor risk assessments, procurement due diligence, or regulatory disclosures should review whether those references remain accurate under the updated policy.

Why it matters

  • ·Regulatory exposure: Organizations that referenced Anthropic's safety pause commitment in regulatory filings or AI governance disclosures may now hold inaccurate representations of their supplier's risk controls, creating potential compliance gaps if regulators scrutinize third-party AI governance documentation.
  • ·Operational impact: Procurement and vendor management workflows that relied on the RSP's precautionary halt mechanism as a supplier-level safety assurance for Claude-based integrations must be updated to reflect the revised framework and any new or absent equivalent controls.
  • ·Organizational risk: The absence of publicly detailed replacement controls in RSP version 3.0 increases uncertainty in third-party AI risk assessments, making it harder for compliance teams to evaluate whether Anthropic's governance commitments continue to meet internal risk acceptance thresholds.

Governance controls affected

What to do now

  • Audit all internal risk documentation, regulatory disclosures, and procurement records that cite Anthropic's Responsible Scaling Policy to identify references to the safety pause commitment and update them to reflect RSP version 3.0.
  • Request updated vendor governance documentation from Anthropic or authorized resellers that details the replacement controls introduced in RSP version 3.0, and evaluate whether those controls satisfy your organization's third-party AI risk acceptance criteria.
  • Review AI vendor contract requirements for Claude-based products to determine whether contractual safety commitments were linked to specific RSP provisions and whether amendments or addenda are now needed.
  • Reassess the AI risk classification for any Claude-dependent systems under HOC-001, particularly where the safety pause mechanism was factored into the original risk tier determination.
  • Update your third-party AI risk assessment templates to include a version-tracking field for supplier governance frameworks, ensuring future policy changes by AI vendors are detected and reviewed systematically.

What to watch next

Compliance teams should monitor Anthropic's official communications and policy publications for a detailed public disclosure of the replacement controls introduced in RSP version 3.0, as the current absence of specifics limits accurate vendor risk assessment. Teams should also track whether frontier AI safety commitments become subject to mandatory disclosure requirements under emerging regulatory frameworks in the EU, UK, or US, which could affect how voluntary policies like the RSP are weighted in compliance evaluations. Any enforcement actions or regulatory guidance referencing voluntary AI governance frameworks at major labs would signal a shift toward greater scrutiny of supplier-level safety documentation.

Stay ahead of stories like this

Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.

Powered by Buttondown.

Related Coverage

Research2026-09-02

Third-Party Frontier AI Auditing Needs Deep Access and Independent Evidence, Report Finds

A research paper from Governance.ai proposes a framework for rigorous third-party auditing of frontier AI developers' safety and security practices. The paper argues that meaningful audits require secure, privileged access to non-public information rather than reliance on developer self-reporting. It has direct implications for enterprise assurance programs that depend on vendor-supplied safety claims.

Enforcement2026-08-29

Sony and Warner Sue Anthropic Over Training Data, Exposing Vendor IP Risk

Sony Music and Warner Chappell have filed a copyright infringement lawsuit against Anthropic in the US District Court for the Northern District of California, alleging that tens of thousands of protected works were used to train Claude without authorization. The complaint seeks up to $150,000 per infringed work and up to $25,000 per instance of stripped copyright metadata, with total exposure potentially reaching several billion dollars. Co-founders Dario Amodei and Benjamin Mann are named as individual defendants.

Corporate Policy2026-08-22

Claude Jailbreak Exposes Content Policy Gap Across Azure and AWS Deployments

TechCrunch testing confirmed that Claude Opus 4.6, Opus 3, and Haiku 4.5 can be manipulated via a multiturn social engineering technique into generating sexually explicit content that Anthropic's universal usage policy explicitly prohibits. The affected models remain available through Anthropic's API and third-party cloud platforms including Azure AI Foundry and Amazon Bedrock. Colorado's recently enacted age-verification law for conversational AI adds a specific regulatory dimension, as the bypass raises questions about whether Anthropic's safeguards meet the statutory 'technically feasible measures' standard.