AI Governance Institute
← News
Research2026-08-04

LLMs Fail on High-Dimensional Tabular Data, Exposing Fitness-for-Purpose Gaps

Source

Why Large Language Models Fail at Tabular Prediction

arXiv / Independent Researchers

What happened

Researchers Marta Garnelo and Wojciech Czarnecki published Why Large Language Models Fail at Tabular Prediction on arXiv, reporting results from controlled experiments across 31 benchmark datasets. The study tested a frontier LLM in a pure inference regime without fine-tuning or external tools, systematically evaluating five candidate explanations for why LLMs underperform classical models on tabular prediction. Of the hypotheses tested, input dimensionality emerged as the decisive factor: as the number of input features grows, LLM accuracy declines, while classical methods hold flat or improve. The finding is directly relevant to enterprise use cases that rely on high-dimensional structured data, including fraud detection, credit risk scoring, anti-money laundering screening, and automated compliance monitoring. Organizations that selected LLMs for these tasks based on general capability benchmarks rather than domain-specific tabular evaluations may be operating with models that are materially less accurate than available alternatives.

Why it matters

  • ·Enterprises that deployed LLMs in fraud detection, AML screening, or risk scoring workflows without high-dimensional tabular benchmarking may be accepting degraded accuracy without knowing it, creating direct regulatory exposure in sectors where model performance standards are enforced.
  • ·Procurement and intake controls built around general capability benchmarks rather than task-specific evaluation are structurally insufficient to catch this failure mode, meaning the gap is as much a process problem as a model problem.
  • ·If classical baselines demonstrably outperform LLMs on the actual data structures used in high-stakes decisions, continued use of LLMs in those roles may be difficult to justify under explainability and fitness-for-purpose obligations increasingly embedded in frameworks like the EU AI Act Implementation Timeline Update and sector-specific guidance from financial regulators.

Governance controls affected

What to do now

  • ☐Audit every active LLM deployment used for tabular prediction tasks (fraud scoring, risk modeling, AML, compliance monitoring) and document the number of input features used in each workflow.
  • ☐Re-run fitness-for-purpose evaluations comparing LLM performance against classical baselines specifically on high-dimensional subsets of your production data, not just general benchmarks used at procurement.
  • ☐Update your model intake and approval workflow (MGV-002) to require tabular-specific benchmark validation for any LLM being considered for structured data prediction tasks.
  • ☐Flag existing deployments where dimensionality was not tested as a risk variable and escalate to model owners for remediation prioritization.
  • ☐Review vendor model cards and documentation for any claims about tabular or structured data performance to determine whether those claims were validated under high-dimensional conditions.

What to watch next

Regulatory guidance on model fitness-for-purpose in financial services is tightening, with the Treasury Department AI Risk Management Framework for Financial Services already setting out voluntary expectations around performance validation in risk-sensitive workflows. Academic findings like this one are increasingly cited in enforcement proceedings and supervisory reviews as evidence of what a reasonable organization should have known at deployment time. Compliance teams should also monitor whether sector regulators in insurance and healthcare issue updated model validation guidance that incorporates structured data performance requirements for LLMs specifically.

Related Coverage

Enforcement2026-09-29

IRS Deployed High-Impact AI With No Testing Records in 80% of Cases

The Treasury Inspector General for Tax Administration (TIGTA) found that four of five high-impact IRS AI use cases had no testing documentation. This was true even though data quality checks were being performed. TIGTA concluded this creates undetectable risk of inaccurate, biased, or unreliable AI outputs. IRS management agreed to standardize procedures and complete AI impact assessments for deployed systems by November 2026.

Research2026-09-26

BIS Warns AI Strains Core Bank Supervisory Expectations on Model Governance

The Bank for International Settlements (BIS) published a speech on September 18, 2026, signaling that advanced AI and large language models (LLMs) are outpacing existing supervisory expectations for banks. The speech identifies governance, model validation, independent review, and explainability as the primary stress points. Banks and their enterprise counterparts in financial services should treat this as a forward signal that supervisors will raise the bar on AI model oversight.

Enforcement2026-09-30

SBA's AI Fraud Pilot Never Classified as High-Impact, OIG Finds

The SBA's Office of Inspector General found that a Palantir-powered AI fraud detection pilot for COVID-19 loan programs was never classified as a high-impact use case under OMB guidance. As a result, required safeguards including impact assessments, human oversight mechanisms, and borrower appeals processes were never put in place. The OIG issued six recommendations, including establishing a formal process for identifying and documenting high-impact AI use cases.