AI Governance Institute
← News

Anthropic's 'Project Panama' Exposes Training Data Sourcing as a Supply-Chain Risk

What happened

Reporting by Ars Technica in Booksellers suspect AI firms are buying and then destroying rare books documents accounts from rare booksellers who flagged a pattern of large, non-negotiated bulk purchases that they now suspect were sourced by AI companies for training data extraction. The reporting draws on a 2025 lawsuit that named Anthropic specifically, alleging that its internal effort known as 'Project Panama' resulted in the physical destruction of millions of print books. Booksellers described telltale signs including orders placed without any price negotiation, a behavior inconsistent with ordinary collectors or institutional buyers. The practice, if confirmed at scale across multiple AI developers, would represent a form of data acquisition that bypasses licensing agreements, obscures provenance records, and leaves no auditable trail for downstream enterprise users. The story adds a physical-world dimension to already active debates about California Generative AI Transparency Requirements - AB 2013 and EU General-Purpose AI Model Training Data Public Summary Template, both of which impose or anticipate disclosure obligations on training data origins.

Why it matters

  • ·Training data provenance is becoming a litigation target. Enterprises deploying models built on undisclosed or improperly obtained training data may face secondary exposure if regulators or courts determine that downstream commercial use amplifies the original violation. Frameworks including the EU General-Purpose AI Model Training Data Public Summary Template require general-purpose AI providers to disclose training data summaries, but those templates cannot catch sourcing practices that were deliberately concealed at acquisition.
  • ·Vendor due diligence programs were not designed to probe physical procurement pipelines. Standard AI vendor assessments focus on model cards, safety evaluations, and data use agreements, none of which would surface a covert book-destruction program. Compliance teams must now consider whether their third-party AI risk assessments include explicit questions about how training data was physically acquired, not merely whether it was licensed.
  • ·The reputational and ESG dimensions are material. Destruction of rare and culturally significant books at scale, if attributed to AI training programs, creates a narrative risk that board-level AI governance committees will need to address proactively. Organizations that have made public commitments to responsible AI sourcing face credibility pressure if their model vendors are linked to practices of this kind.

Governance controls affected

What to do now

  • Add explicit questions to your AI vendor due diligence questionnaire covering physical training data acquisition methods, including whether the vendor has purchased, destroyed, or physically processed copyrighted materials to extract training data.
  • Review existing AI vendor contracts to determine whether training data sourcing representations and warranties extend to physical procurement practices, and negotiate addenda where they do not.
  • Assess whether any foundation models currently in use or under evaluation are named in the 2025 lawsuit or related proceedings, and flag those vendors for enhanced monitoring under your third-party risk program.
  • Brief your board AI governance committee or risk committee on the reputational and ESG dimensions of training data sourcing practices, particularly where enterprise AI commitments reference responsible or ethical AI use.
  • Map your organization's reliance on general-purpose AI models against applicable training data disclosure requirements, including those arising under California AB 2013 and the EU GPAI training data summary template, and identify disclosure gaps that a covert sourcing program would create.

What to watch next

Active litigation against Anthropic over 'Project Panama' is likely to produce additional discovery that further details the scale, methods, and internal decision-making behind covert book-acquisition programs. Compliance teams should monitor court filings for any findings that implicate other AI developers, since bookseller accounts suggest Anthropic may not be alone. Regulatory attention to training data transparency under California Generative AI Transparency Requirements - AB 2013 and the EU's GPAI disclosure framework is already intensifying, and enforcement agencies may treat concealed physical sourcing as an aggravating factor. The H.R.8094 - AI Foundation Model Transparency Act of 2026 is also advancing in Congress and could impose federal-level training data disclosure requirements that make these sourcing questions a statutory compliance obligation rather than a contractual one.

Stay ahead of stories like this

Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.

Powered by Buttondown.

Related Coverage

Corporate Policy2026-08-31

Redacted Anthropic Risk Report on Claude Mythos Preview Leaves Compliance Teams Without a Safety Case

Anthropic published a formal risk report in August 2026 referencing Claude Mythos Preview, a model available through its limited-access Glasswing program. The report signals a safety-review posture but is substantially redacted, leaving enterprise buyers without the full evaluation findings needed to assess suitability for regulated deployment. Compliance teams should not treat report existence as a substitute for complete model documentation.

Enforcement2026-08-29

Sony and Warner Sue Anthropic Over Training Data, Exposing Vendor IP Risk

Sony Music and Warner Chappell have filed a copyright infringement lawsuit against Anthropic in the US District Court for the Northern District of California, alleging that tens of thousands of protected works were used to train Claude without authorization. The complaint seeks up to $150,000 per infringed work and up to $25,000 per instance of stripped copyright metadata, with total exposure potentially reaching several billion dollars. Co-founders Dario Amodei and Benjamin Mann are named as individual defendants.

Enforcement2026-09-01

EFF Fights 'Market Dilution' Theory That Would End Fair Use for AI Training

The Electronic Frontier Foundation has filed amicus briefs in Concord Music Group v. Anthropic and In re Mosaic LLM Litigation, urging courts to reject a copyright liability theory that would allow rightsholders to block AI training on any work that competes with their existing markets. The EFF argues that accepting this 'market dilution' theory would effectively gut fair use doctrine as a permissible basis for training data ingestion. Enterprise compliance teams whose training data programs rely on fair use as a legal foundation should treat both cases as active, high-priority litigation risk.