AI Governance Institute
← News
Research2026-07-31

ITU Report Sets Global Baseline for Pre-Deployment Testing and Watermarking Norms

What happened

The ITU released the Annual AI Governance Report 2025, a global synthesis of emerging AI governance norms covering pre-deployment registration, licensing proposals, mandatory safety testing, red-teaming requirements, and post-deployment transparency obligations such as content watermarking. The report is non-binding but carries significant weight as a credible multilateral reference point that regulators and standards bodies across jurisdictions are likely to cite. It arrives as multiple frameworks, including the EU AI Act Implementation Timeline, the OECD AI Principles, and UNESCO Recommendation on the Ethics of Artificial Intelligence, are moving pre-deployment testing and provenance disclosure from voluntary commitments toward enforceable requirements. For enterprises, the report consolidates three governance pressure points into a single reference: the adequacy of model release gates, the evidentiary quality of safety evaluations, and the maturity of content provenance controls. The publication follows growing scrutiny of whether formal pre-deployment testing norms translate into enforceable audit evidence, a gap highlighted by recent case studies exposing documentation and accountability gaps at AI release decision points.

Why it matters

  • ·The report signals that pre-deployment testing and red-teaming are transitioning from best-practice guidance to anticipated regulatory requirements. Compliance programs that treat safety evaluations as informal or ad hoc reviews now face increasing exposure as regulators in the EU, the UK, and several US states look to frameworks like this one to define what adequate pre-deployment controls look like.
  • ·Watermarking and content provenance requirements are explicitly framed in the report as post-deployment transparency obligations, not optional features. Enterprises that have not yet assessed their readiness to label and track AI-generated content risk being underprepared for mandatory disclosure regimes emerging under instruments such as the EU Code of Practice on Marking and Labelling of AI-Generated Content and parallel measures in several jurisdictions.
  • ·The report's treatment of model registration and licensing proposals elevates the governance risk for organizations that deploy models without formal intake and approval records. Regulatory bodies examining enterprise AI programs will increasingly expect documented approval gates, audit trails from safety evaluations, and evidence that red-teaming results informed deployment decisions, not merely that testing occurred.

Governance controls affected

What to do now

  • ☐Map your current pre-deployment testing documentation against the report's red-teaming and safety evaluation norms to identify gaps that could become audit findings under emerging mandatory regimes.
  • ☐Review your model release gate process to confirm it produces audit-ready evidence of safety evaluations, not just internal sign-off records, before models reach production.
  • ☐Assess whether your AI-generated content workflows support watermarking or provenance labeling at scale, and identify tools or process changes needed to meet anticipated mandatory disclosure requirements.
  • ☐Update your multi-jurisdiction compliance mapping to include the ITU report as a reference document alongside binding frameworks, so your regulatory horizon scanning captures where convergent norms are forming.
  • ☐Brief your board or AI governance committee on the report's registration and licensing proposals as forward-looking regulatory signals that may affect your model intake and vendor management programs within the next one to two years.

What to watch next

Compliance teams should monitor whether the ITU report's pre-deployment testing norms are cited in upcoming enforcement guidance or rulemaking by the EU AI Office, national competent authorities under the EU AI Act Implementation Timeline, or US federal agencies developing sector-specific AI requirements. The convergence between this report and the Singapore Consensus on Global AI Safety Research Priorities suggests that multilateral alignment on mandatory safety testing is accelerating. Organizations in regulated sectors, particularly financial services, healthcare, and critical infrastructure, should treat 2026 as the window to formalize their pre-deployment evaluation programs before voluntary norms harden into enforceable obligations.

Stay ahead of stories like this

Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.

Powered by Buttondown.

Related Coverage

Research2026-09-23

Embassy-Amplified BRICS Deepfake Exposes the Pre-Publication Verification Gap

An AI-generated image falsely depicting world leaders at the BRICS summit was shared by embassy social media accounts before fact-checking caught it. Resemble AI's Deepfake Watchlist documented the incident as a case study in how synthetic media gains institutional credibility through official amplification. The incident reveals that watermark awareness does not translate into actual provenance checks before content is published.

Research2026-09-17

SynthID-Text Watermarking Weakens Safety Guardrails, Lasso Security Finds

Lasso Security researcher Andrea Siposova found that SynthID-Text watermarking alters how LLMs respond to harmful prompts, including bypassing safety refusals. The effect, which Siposova calls 'sampling drift,' extends into agentic pipelines by influencing which tools agents invoke. Anthropic has committed to deploying SynthID-Text in future Claude models, partly in response to EU AI Act provenance requirements.

Research2026-09-26

UK AISI Study: AI Shifted Political Views by 10 Points in 42,000-Person Trial

Research with UK AI Security Institute support tested AI language models on over 42,000 participants. It found average political and attitudinal shifts of roughly 10 percentage points. AI-generated persuasive content outperformed static social media messages by 41-52%. Models fine-tuned specifically to maximize persuasion produced inaccurate claims in nearly a third of responses.