AI Governance Institute
← News

OpenAI Cannot Rule Out Training on Researchers' Codex Sessions

What happened

OpenAI announced a claimed solution to the Navier-Stokes Millennium Prize Problem, reportedly developed using an internal AI model running 10,000 concurrent agents, as reported by Drama swirls around OpenAI's legendary mathematical milestone. The announcement drew immediate challenge from NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpoge, who had published related findings one day earlier and expressed concern that OpenAI's model may have reached a parallel proof route by drawing on data from their Codex sessions. OpenAI issued a statement denying that any specific user data was accessed, but it explicitly acknowledged that it could not rule out the possibility that de-identified training data derived from Buckmaster and Alpoge's Codex usage had influenced the model's approach. That partial denial is the governance event: a major AI developer has confirmed in public that its assurance about user data stops short of covering de-identified aggregates used in training. The disclosure arrives in a research context but the implication extends to every enterprise customer using OpenAI's developer tools for proprietary, competitive, or IP-sensitive work.

Why it matters

  • ·Standard vendor representations that specific session data is not accessed do not address whether de-identified aggregates of user inputs are used in model training, leaving a gap that most enterprise vendor contracts and data classification programs have not addressed.
  • ·Organizations conducting proprietary research, competitive analysis, or trade-secret-adjacent work through AI coding tools like Codex now face an unresolved question about whether that work could shape future model behavior in ways that benefit competitors or third parties who use the same platform.
  • ·Procurement and legal teams relying on data processing agreements to govern AI vendor data use should treat this disclosure as a signal that contractual scope needs explicit language covering de-identified training use, not just identified session-level data access, to meet obligations under frameworks like the NIST Artificial Intelligence Risk Management Framework Playbook.

Governance controls affected

What to do now

  • Review your current data processing agreements and terms of service with OpenAI and comparable AI developer tool vendors to determine whether de-identified training data use is explicitly addressed or excluded.
  • Update your data classification policy to include a category for inputs to AI coding and research tools that carry proprietary, competitive, or IP-sensitive content, with handling rules that account for training data risk.
  • Require legal and procurement teams to add explicit contractual language covering de-identified aggregate training use when renewing or entering AI developer tool contracts.
  • Conduct a retroactive review of which teams have used Codex or equivalent tools for work involving trade secrets, pending patents, unpublished research, or competitive strategy, and assess whether that use falls within current data governance controls.
  • Brief your AI governance committee and general counsel on this disclosure as a precedent that vendor assurances about 'specific session data' are not equivalent to assurances about training data exclusion.

What to watch next

Compliance teams should monitor whether OpenAI or other major AI developer tool vendors revise their data use policies to address de-identified training use explicitly, as this disclosure is likely to attract scrutiny from privacy regulators in jurisdictions where data minimization and purpose-limitation requirements apply. Regulatory bodies in the EU may treat de-identified training use as a distinct processing purpose requiring separate legal basis under applicable data protection frameworks. The controversy also signals that IP and trade secret litigation involving AI training data is likely to expand beyond the creative industry cases already in litigation, and enterprises should watch for guidance from bar associations and IP counsel on how this affects confidentiality obligations for legal and research workflows conducted through AI tools.

Stay ahead of stories like this

Get every US AI governance development like this one, plus the rest of the week's developments. Every Thursday.

Powered by Buttondown.

Related Coverage

Corporate Policy2026-09-03

Meta's 95% API Discount Creates a Data Classification Forcing Function

Meta is offering enterprise customers roughly a 95% reduction in Muse Spark API costs in exchange for consent to use their prompts and model outputs as training data. The structure creates a direct financial incentive to share workflow data with a model provider, raising compliance questions about which data enterprises can lawfully contribute. Organizations without a mature data classification policy face meaningful exposure before they can make an informed procurement decision.

Enforcement2026-08-27

Grok CSAM Lawsuit Sets a Training Data Provenance Liability Benchmark

A federal lawsuit filed by a child sex abuse material survivor alleges that xAI trained its Grok models on CSAM identified via hash lists maintained by NCMEC and the Canadian Centre for Child Protection. The complaint also alleges that xAI's terms of service create a training pipeline that recycles public posts and model outputs without explicit exclusion categories for illegal content. Enterprise compliance teams now have a concrete litigation template against which to audit their own training data provenance and vendor due diligence controls.

Corporate Policy2026-09-02

Mistral Default Opt-In for Training Data Creates GDPR Exposure on Non-Enterprise Tiers

Mistral AI has updated its data policy so that user conversations and uploaded documents are used for model training by default on its Vibe consumer and standard API tiers. Enterprise customers on the Vibe Enterprise plan are opted out by default, with admin-level controls to manage opt-in. The split configuration requirement between Vibe and API surfaces creates separate compliance exposure that must be managed independently.