OpenAI Cannot Rule Out Training on Researchers' Codex Sessions
What happened
OpenAI announced a claimed solution to the Navier-Stokes Millennium Prize Problem, reportedly developed using an internal AI model running 10,000 concurrent agents, as reported by Drama swirls around OpenAI's legendary mathematical milestone. The announcement drew immediate challenge from NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpoge, who had published related findings one day earlier and expressed concern that OpenAI's model may have reached a parallel proof route by drawing on data from their Codex sessions. OpenAI issued a statement denying that any specific user data was accessed, but it explicitly acknowledged that it could not rule out the possibility that de-identified training data derived from Buckmaster and Alpoge's Codex usage had influenced the model's approach. That partial denial is the governance event: a major AI developer has confirmed in public that its assurance about user data stops short of covering de-identified aggregates used in training. The disclosure arrives in a research context but the implication extends to every enterprise customer using OpenAI's developer tools for proprietary, competitive, or IP-sensitive work.
Why it matters
- ·Standard vendor representations that specific session data is not accessed do not address whether de-identified aggregates of user inputs are used in model training, leaving a gap that most enterprise vendor contracts and data classification programs have not addressed.
- ·Organizations conducting proprietary research, competitive analysis, or trade-secret-adjacent work through AI coding tools like Codex now face an unresolved question about whether that work could shape future model behavior in ways that benefit competitors or third parties who use the same platform.
- ·Procurement and legal teams relying on data processing agreements to govern AI vendor data use should treat this disclosure as a signal that contractual scope needs explicit language covering de-identified training use, not just identified session-level data access, to meet obligations under frameworks like the NIST Artificial Intelligence Risk Management Framework Playbook.
Governance controls affected
What to do now
- ☐Review your current data processing agreements and terms of service with OpenAI and comparable AI developer tool vendors to determine whether de-identified training data use is explicitly addressed or excluded.
- ☐Update your data classification policy to include a category for inputs to AI coding and research tools that carry proprietary, competitive, or IP-sensitive content, with handling rules that account for training data risk.
- ☐Require legal and procurement teams to add explicit contractual language covering de-identified aggregate training use when renewing or entering AI developer tool contracts.
- ☐Conduct a retroactive review of which teams have used Codex or equivalent tools for work involving trade secrets, pending patents, unpublished research, or competitive strategy, and assess whether that use falls within current data governance controls.
- ☐Brief your AI governance committee and general counsel on this disclosure as a precedent that vendor assurances about 'specific session data' are not equivalent to assurances about training data exclusion.
What to watch next
Compliance teams should monitor whether OpenAI or other major AI developer tool vendors revise their data use policies to address de-identified training use explicitly, as this disclosure is likely to attract scrutiny from privacy regulators in jurisdictions where data minimization and purpose-limitation requirements apply. Regulatory bodies in the EU may treat de-identified training use as a distinct processing purpose requiring separate legal basis under applicable data protection frameworks. The controversy also signals that IP and trade secret litigation involving AI training data is likely to expand beyond the creative industry cases already in litigation, and enterprises should watch for guidance from bar associations and IP counsel on how this affects confidentiality obligations for legal and research workflows conducted through AI tools.
Stay ahead of stories like this
Get every US AI governance development like this one, plus the rest of the week's developments. Every Thursday.
