AI Governance Institute
← News
Research2026-04-19

AI Incidents Rising and Responsible AI Evals Still Rare Among Major Developers, Stanford HAI 2025 Index Finds

What happened

Stanford University's Human-Centered Artificial Intelligence institute published the 2025 AI Index Report on April 1, 2025, offering a comprehensive global analysis of AI research, development, and governance trends. The report documents a measurable increase in AI-related incidents and finds that standardized responsible AI evaluations remain uncommon among major industrial model developers, identifying a persistent gap between organizational acknowledgment of responsible AI risks and concrete remediation steps. Emerging benchmarks including HELM Safety, AIR-Bench, and FACTS are highlighted as tools designed to assess model safety and factuality, though adoption across the industry remains limited. The report also notes accelerated regulatory output across multiple jurisdictions, citing frameworks from the OECD, European Union, and United Nations that increasingly emphasize transparency and trustworthiness requirements for AI systems. Organizations operating in EU jurisdictions face near-term deadlines under the EU AI Act that make alignment with these frameworks particularly urgent.

Why it matters

  • ·Regulators across the EU, OECD, and UN are transitioning from principle-setting toward enforceable standards, meaning organizations that lack documented responsible AI evaluation processes face growing legal and audit exposure.
  • ·The documented rise in AI incidents combined with the scarcity of standardized RAI evaluations signals that voluntary commitments have not consistently translated into operational practice, increasing the likelihood that regulators will mandate demonstrable safety evidence rather than accepting policy statements.
  • ·Organizations that have not yet adopted structured benchmarking tools such as HELM Safety or AIR-Bench risk being unprepared for auditor and regulatory scrutiny that increasingly expects systematic, evidence-based governance rather than ad hoc or informal approaches.

Governance controls affected

What to do now

  • ☐Assess the maturity of current responsible AI evaluation processes against named benchmarks including HELM Safety, AIR-Bench, and FACTS, and document gaps relative to internal governance standards.
  • ☐Map existing AI governance documentation against OECD and EU AI Act transparency and trustworthiness requirements, prioritizing controls relevant to high-risk system classifications.
  • ☐Establish or strengthen incident tracking mechanisms capable of capturing, classifying, and escalating AI-related failures in a format that satisfies anticipated regulatory disclosure and audit expectations.
  • ☐Review and update the AI Incident Response Playbook to reflect the increased incident volume documented in the report and align severity classification with emerging regulatory disclosure thresholds.
  • ☐Initiate a gap analysis comparing current model evaluation and pre-production approval processes against structured safety and factuality benchmarks to identify areas requiring remediation before near-term EU AI Act deadlines.

What to watch next

Compliance teams should monitor the progression of EU AI Act implementation deadlines, as obligations for high-risk AI systems are approaching enforcement phases that will require documented evidence of safety evaluations rather than policy-level commitments. Ongoing guidance from the OECD and United Nations on AI transparency requirements should be tracked for signals of convergence toward harmonized global standards that could affect organizations across multiple jurisdictions. The adoption trajectory of benchmarks such as HELM Safety and AIR-Bench should also be monitored, as broader industry uptake may lead regulators and auditors to treat these tools as de facto compliance reference points.

Stay ahead of stories like this

Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.

Powered by Buttondown.

Related Coverage

Research2026-09-28

Six-Pillar AI Governance Model Sets Enterprise Program Maturity Benchmark

Concurrency, a technology consulting firm, has published a practitioner framework organizing enterprise AI governance into six pillars: inventory, validation, monitoring, explainability, fairness testing, and incident response. The framework targets enterprises that have deployed AI but lack structured approval gates, continuous monitoring, or audit evidence. It provides a replicable operating model that compliance teams can use to measure and close program gaps.

Enforcement2026-09-28

EU AI Office Inspections Target Hiring, Credit, and Healthcare AI

The European AI Office and national market surveillance authorities launched coordinated compliance inspections of high-risk AI systems in September 2026. The inspections focus on resume-screening tools, credit-assessment systems, and healthcare triage applications. Organizations lacking documentation, audit trails, and rapid remediation plans are the primary targets.

Corporate Policy2026-09-26

HUD Deploys AI Grants Monitor by September 30 Without Finalized Governance Rules

The U.S. Department of Housing and Urban Development (HUD) is preparing to launch an AI-powered grants monitoring system called HUGS by September 30, 2026. Built under a $500,000 contract with Palantir, the system will analyze every vendor payment made by grant recipients and flag outliers for human review. Governance procedures for the system have not yet been finalized, and housing and technology experts have raised concerns about disparate impact on smaller grantees.