LLMs Fail on High-Dimensional Tabular Data, Exposing Fitness-for-Purpose Gaps
What happened
Researchers Marta Garnelo and Wojciech Czarnecki published Why Large Language Models Fail at Tabular Prediction on arXiv, reporting results from controlled experiments across 31 benchmark datasets. The study tested a frontier LLM in a pure inference regime without fine-tuning or external tools, systematically evaluating five candidate explanations for why LLMs underperform classical models on tabular prediction. Of the hypotheses tested, input dimensionality emerged as the decisive factor: as the number of input features grows, LLM accuracy declines, while classical methods hold flat or improve. The finding is directly relevant to enterprise use cases that rely on high-dimensional structured data, including fraud detection, credit risk scoring, anti-money laundering screening, and automated compliance monitoring. Organizations that selected LLMs for these tasks based on general capability benchmarks rather than domain-specific tabular evaluations may be operating with models that are materially less accurate than available alternatives.
Why it matters
- ·Enterprises that deployed LLMs in fraud detection, AML screening, or risk scoring workflows without high-dimensional tabular benchmarking may be accepting degraded accuracy without knowing it, creating direct regulatory exposure in sectors where model performance standards are enforced.
- ·Procurement and intake controls built around general capability benchmarks rather than task-specific evaluation are structurally insufficient to catch this failure mode, meaning the gap is as much a process problem as a model problem.
- ·If classical baselines demonstrably outperform LLMs on the actual data structures used in high-stakes decisions, continued use of LLMs in those roles may be difficult to justify under explainability and fitness-for-purpose obligations increasingly embedded in frameworks like the EU AI Act Implementation Timeline Update and sector-specific guidance from financial regulators.
Governance controls affected
What to do now
- ☐Audit every active LLM deployment used for tabular prediction tasks (fraud scoring, risk modeling, AML, compliance monitoring) and document the number of input features used in each workflow.
- ☐Re-run fitness-for-purpose evaluations comparing LLM performance against classical baselines specifically on high-dimensional subsets of your production data, not just general benchmarks used at procurement.
- ☐Update your model intake and approval workflow (MGV-002) to require tabular-specific benchmark validation for any LLM being considered for structured data prediction tasks.
- ☐Flag existing deployments where dimensionality was not tested as a risk variable and escalate to model owners for remediation prioritization.
- ☐Review vendor model cards and documentation for any claims about tabular or structured data performance to determine whether those claims were validated under high-dimensional conditions.
What to watch next
Regulatory guidance on model fitness-for-purpose in financial services is tightening, with the Treasury Department AI Risk Management Framework for Financial Services already establishing expectations around performance validation in risk-sensitive workflows. Academic findings like this one are increasingly cited in enforcement proceedings and supervisory reviews as evidence of what a reasonable organization should have known at deployment time. Compliance teams should also monitor whether sector regulators in insurance and healthcare issue updated model validation guidance that incorporates structured data performance requirements for LLMs specifically.
Stay ahead of stories like this
Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.
