AI Governance Institute logo
AI Governance Institute

Intelligence for Compliance and GRC Teams

← News
Research2026-08-04

LLMs Fail on High-Dimensional Tabular Data, Exposing Fitness-for-Purpose Gaps

Source

Why Large Language Models Fail at Tabular Prediction

arXiv / Independent Researchers

What happened

Researchers Marta Garnelo and Wojciech Czarnecki published Why Large Language Models Fail at Tabular Prediction on arXiv, reporting results from controlled experiments across 31 benchmark datasets. The study tested a frontier LLM in a pure inference regime without fine-tuning or external tools, systematically evaluating five candidate explanations for why LLMs underperform classical models on tabular prediction. Of the hypotheses tested, input dimensionality emerged as the decisive factor: as the number of input features grows, LLM accuracy declines, while classical methods hold flat or improve. The finding is directly relevant to enterprise use cases that rely on high-dimensional structured data, including fraud detection, credit risk scoring, anti-money laundering screening, and automated compliance monitoring. Organizations that selected LLMs for these tasks based on general capability benchmarks rather than domain-specific tabular evaluations may be operating with models that are materially less accurate than available alternatives.

Why it matters

  • ·Enterprises that deployed LLMs in fraud detection, AML screening, or risk scoring workflows without high-dimensional tabular benchmarking may be accepting degraded accuracy without knowing it, creating direct regulatory exposure in sectors where model performance standards are enforced.
  • ·Procurement and intake controls built around general capability benchmarks rather than task-specific evaluation are structurally insufficient to catch this failure mode, meaning the gap is as much a process problem as a model problem.
  • ·If classical baselines demonstrably outperform LLMs on the actual data structures used in high-stakes decisions, continued use of LLMs in those roles may be difficult to justify under explainability and fitness-for-purpose obligations increasingly embedded in frameworks like the EU AI Act Implementation Timeline Update and sector-specific guidance from financial regulators.

Governance controls affected

What to do now

  • Audit every active LLM deployment used for tabular prediction tasks (fraud scoring, risk modeling, AML, compliance monitoring) and document the number of input features used in each workflow.
  • Re-run fitness-for-purpose evaluations comparing LLM performance against classical baselines specifically on high-dimensional subsets of your production data, not just general benchmarks used at procurement.
  • Update your model intake and approval workflow (MGV-002) to require tabular-specific benchmark validation for any LLM being considered for structured data prediction tasks.
  • Flag existing deployments where dimensionality was not tested as a risk variable and escalate to model owners for remediation prioritization.
  • Review vendor model cards and documentation for any claims about tabular or structured data performance to determine whether those claims were validated under high-dimensional conditions.

What to watch next

Regulatory guidance on model fitness-for-purpose in financial services is tightening, with the Treasury Department AI Risk Management Framework for Financial Services already establishing expectations around performance validation in risk-sensitive workflows. Academic findings like this one are increasingly cited in enforcement proceedings and supervisory reviews as evidence of what a reasonable organization should have known at deployment time. Compliance teams should also monitor whether sector regulators in insurance and healthcare issue updated model validation guidance that incorporates structured data performance requirements for LLMs specifically.

Stay ahead of stories like this

Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.

Powered by Buttondown.

Related Coverage

Research2026-08-17

Keyrus 2026 Guide Sets a Baseline Operating Model for AI Governance Programs

Consulting firm Keyrus has published a practitioner guide outlining how enterprises should structure AI governance programs in 2026, emphasizing four foundational elements: a complete AI inventory, risk-based prioritization, cross-functional governance teams, and oversight of vendor-supplied models. The guide provides a replicable operating model that compliance teams can adapt and pair with existing controls. It targets organizations at any stage of AI governance maturity.

Research2026-08-17

KPMG Frames AI Governance as a Model Risk Problem, Not a Separate Silo

KPMG has published a guide positioning AI oversight as an extension of existing model risk management structures rather than a standalone governance program. The guide organizes AI oversight around four pillars: governance, development, validation, and monitoring. Compliance teams are advised to integrate AI controls into familiar model risk frameworks rather than build parallel processes.

Research2026-08-19

EU AI Act Enforcement Has Begun: Documentation Gaps Now Draw Regulator Attention

The Future of Life Institute's EU AI Act Newsletter #108 reports that enforcement activity under the EU AI Act is now underway, shifting the regulation from a planning horizon to an active compliance obligation. The newsletter tracks emerging enforcement patterns and flags documentation and transparency obligations as the most immediate areas of exposure. Compliance teams operating in EU-regulated markets should use enforcement signals to stress-test existing control mappings and update their conformity assessment processes.