AI Governance Institute
← News
Research2026-09-02

Third-Party Frontier AI Auditing Needs Deep Access and Independent Evidence, Report Finds

What happened

Governance.ai published Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies, a research paper arguing that current auditing of frontier AI developers is structurally inadequate. The paper contends that auditors need deep, secure access to non-public internal processes, evaluation data, and safety records to form reliable independent opinions about whether a developer's safety and security claims hold up. Without this access, third-party reviews function more as document reviews than genuine assurance exercises. The paper is particularly relevant now that enterprises are relying on published safety frameworks and commitments from developers such as Anthropic and OpenAI as inputs to their own procurement risk assessments. This follows a broader pattern of concern about the reliability of developer-supplied safety narratives, including the finding that vendor AI usage reports systematically filter harmful behavior, and industry discussions prompted when 12 frontier developers published formal AI safety frameworks without a shared standard for verifying their claims.

Why it matters

  • ·Enterprise procurement and vendor risk teams that rely on developer safety cards, system cards, or voluntary commitment documents as compliance evidence now have a credible research basis for questioning whether that evidence is independently verifiable. Regulators in multiple jurisdictions, including the EU under the EU AI Act: AI Literacy and Prohibited AI Systems Provisions (Applicable 2 February 2026), are moving toward requiring substantiated safety assurance rather than self-attestation, which raises the stakes for this gap.
  • ·If regulators or auditors begin to require evidence that vendor safety claims are externally verified, enterprises that have accepted those claims at face value will need to retrofit their vendor due diligence programs. Organizations in financial services, healthcare, and critical infrastructure face the sharpest exposure because sector-specific frameworks in those industries already assume model risk can be independently assessed.
  • ·The paper exposes a structural weakness in how safety assurance is currently scoped: most audit rights in enterprise AI vendor contracts do not extend to the pre-deployment safety evaluation records and red-team findings that a rigorous audit would require. This creates a documentation gap that could become a liability issue when incidents occur and regulators ask what assurance steps were taken before deployment.

Governance controls affected

What to do now

  • Review existing AI vendor contracts to determine whether audit rights provisions allow access to pre-deployment safety evaluation records, red-team findings, and incident logs, and flag contracts that limit review to documentation supplied by the vendor.
  • Map any frontier AI vendor safety claims currently used as procurement evidence against the Governance.ai paper's proposed audit criteria, and document where those claims are supported by independently verifiable data versus self-attestation.
  • Update vendor due diligence questionnaires to require frontier AI developers to disclose whether their safety evaluations have been reviewed by a third party with privileged internal access, and what that review covered.
  • Escalate to the AI governance committee any procurement decisions where a frontier model's safety case rests primarily on the developer's own published materials, particularly for high-risk use cases.
  • Monitor regulatory developments in the EU and US for mandatory third-party audit requirements for frontier AI developers, and prepare a contingency plan for adjusting vendor assurance standards if those requirements become binding.

What to watch next

Compliance teams should track whether the EU AI Office references independent audit access requirements in forthcoming guidance under the European Commission Enforcement Powers for Advanced AI Models under the AI Act, which could make the access standards proposed in this paper a de facto regulatory baseline. Voluntary frameworks negotiated with frontier developers by the US government may also evolve to incorporate structured third-party review requirements, particularly if the White House AI vulnerability-sharing initiative produces audit-related commitments. Watch for whether frontier developers begin amending their terms of service or safety commitments to accommodate auditor access, which would signal that market pressure is beginning to close the gap the paper identifies.

Stay ahead of stories like this

Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.

Powered by Buttondown.

Related Coverage

Research2026-09-02

FLI Safety Index Ranks Frontier AI Firms, Creating a Vendor Benchmarking Obligation

The Future of Life Institute published its AI Safety Index Summer 2026 on August 26, 2026, ranking major frontier AI developers on safety practices and transparency. Anthropic leads across most domains in the ranking. The index gives enterprise compliance teams an external benchmark to use in vendor due diligence, procurement risk assessments, and board-level AI risk reporting.

Enforcement2026-08-31

ChatGPT Designated a Very Large Online Platform Under EU DSA

The European Commission has designated ChatGPT as a Very Large Online Search Engine under the EU Digital Services Act, imposing elevated compliance obligations on OpenAI with a December 2026 deadline. Requirements include protecting minors, curbing illegal content, restricting behavioral advertising, and providing algorithmic transparency. Enterprise deployers using ChatGPT in the EU now face downstream vendor governance obligations tied to this designation.

Enforcement2026-09-02

Alabama AG Subpoena Puts OpenAI Agent Oversight Controls Under State Enforcement Scrutiny

Alabama's attorney general has opened a formal, subpoena-driven investigation into OpenAI and Sam Altman over the company's handling of an agent autonomy incident and its broader oversight practices. The inquiry centers on whether OpenAI's safety review, logging, and third-party impact controls were adequate to prevent or fully explain the agent behavior. The action marks the first known state-level enforcement effort targeting an AI developer's internal governance controls.