AI Governance Institute logo
AI Governance Institute

Intelligence for Compliance and GRC Teams

← News
Research2026-04-19

AI Incidents Rising and Responsible AI Evals Still Rare Among Major Developers, Stanford HAI 2025 Index Finds

What happened

Stanford University's Human-Centered Artificial Intelligence institute published the 2025 AI Index Report on April 1, 2025, offering a comprehensive global analysis of AI research, development, and governance trends. The report documents a measurable increase in AI-related incidents and finds that standardized responsible AI evaluations remain uncommon among major industrial model developers, identifying a persistent gap between organizational acknowledgment of responsible AI risks and concrete remediation steps. Emerging benchmarks including HELM Safety, AIR-Bench, and FACTS are highlighted as tools designed to assess model safety and factuality, though adoption across the industry remains limited. The report also notes accelerated regulatory output across multiple jurisdictions, citing frameworks from the OECD, European Union, and United Nations that increasingly emphasize transparency and trustworthiness requirements for AI systems. Organizations operating in EU jurisdictions face near-term deadlines under the EU AI Act that make alignment with these frameworks particularly urgent.

Why it matters

  • ·Regulators across the EU, OECD, and UN are transitioning from principle-setting toward enforceable standards, meaning organizations that lack documented responsible AI evaluation processes face growing legal and audit exposure.
  • ·The documented rise in AI incidents combined with the scarcity of standardized RAI evaluations signals that voluntary commitments have not consistently translated into operational practice, increasing the likelihood that regulators will mandate demonstrable safety evidence rather than accepting policy statements.
  • ·Organizations that have not yet adopted structured benchmarking tools such as HELM Safety or AIR-Bench risk being unprepared for auditor and regulatory scrutiny that increasingly expects systematic, evidence-based governance rather than ad hoc or informal approaches.

Governance controls affected

What to do now

  • Assess the maturity of current responsible AI evaluation processes against named benchmarks including HELM Safety, AIR-Bench, and FACTS, and document gaps relative to internal governance standards.
  • Map existing AI governance documentation against OECD and EU AI Act transparency and trustworthiness requirements, prioritizing controls relevant to high-risk system classifications.
  • Establish or strengthen incident tracking mechanisms capable of capturing, classifying, and escalating AI-related failures in a format that satisfies anticipated regulatory disclosure and audit expectations.
  • Review and update the AI Incident Response Playbook to reflect the increased incident volume documented in the report and align severity classification with emerging regulatory disclosure thresholds.
  • Initiate a gap analysis comparing current model evaluation and pre-production approval processes against structured safety and factuality benchmarks to identify areas requiring remediation before near-term EU AI Act deadlines.

What to watch next

Compliance teams should monitor the progression of EU AI Act implementation deadlines, as obligations for high-risk AI systems are approaching enforcement phases that will require documented evidence of safety evaluations rather than policy-level commitments. Ongoing guidance from the OECD and United Nations on AI transparency requirements should be tracked for signals of convergence toward harmonized global standards that could affect organizations across multiple jurisdictions. The adoption trajectory of benchmarks such as HELM Safety and AIR-Bench should also be monitored, as broader industry uptake may lead regulators and auditors to treat these tools as de facto compliance reference points.

Stay ahead of stories like this

Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.

Powered by Buttondown.

Related Coverage

Research2026-07-24

S&P Global Identifies Five Governance Principles That Should Anchor Every Enterprise AI Risk Program

S&P Global has published a research report titled 'The AI Governance Challenge' identifying transparency, fairness, privacy, adaptability, and accountability as the five core principles that should structure enterprise AI governance programs. The report is addressed to enterprise risk and compliance leaders and offers design guidance for documentation standards, bias review processes, privacy impact assessments, and accountability structures. It carries no regulatory force but reflects an emerging market consensus from a recognized financial intelligence institution.

Research2026-07-23

Google's ATLAS Study Puts Empirical Numbers on Workforce AI Adoption, Creating New Obligations for Impact Assessments and Transparency Disclosures

Google has published the ATLAS study, a large-scale analysis of 15 million de-identified AI interactions drawn from Gemini App, AI Mode, and the Gemini API. The study finds that while AI touches 68% of occupations, it covers only about 21% of tasks within a typical job, and fewer than 10% of interactions fully automate a task. The findings provide the first major empirical baseline for workforce impact assessments required under an expanding set of AI governance frameworks.

Research2026-07-21

32% of Recent arXiv Papers Flag as AI-Written, Exposing Reliability Gaps in Detection Tools Enterprises Rely On

A study by Unsloth scored 12,750 arXiv preprints and found approximately 32% of recent submissions display markers of machine-generated text, up from a pre-ChatGPT baseline of 0.4%. The prevalence varies sharply by discipline, with computer science papers flagging at 65% and mathematics near 0.7%. The study also documents significant limitations in current AI detection methodology, including blind spots and the inability to distinguish lightly edited from wholly generated text.