Tabular foundation models interpolate physics but cannot model it
Tabular foundation models are outperforming traditional baselines on physics-informed datasets, but a new evaluation reveals a critical limitation: they cannot represent deterministic systems or physical units. Researchers tested four leading TFMs against 316 physics equations and found that while these models excel at interpolation within their training distribution, they lack the structural priors needed to function as genuine physics simulators. This gap matters for scientific computing and industrial applications where unit consistency and noiseless mechanisms are non-negotiable, suggesting that scaling tabular models alone won't bridge the gap between statistical pattern matching and causal physical reasoning.
Modelwire context
ExplainerThe critical finding isn't just that tabular foundation models fail on physics tasks, but that they fail specifically on deterministic (noiseless) systems where there is only one correct answer. This reveals the models are learning statistical associations rather than structural rules, a distinction that benchmarks measuring noisy-data performance would miss entirely.
This connects directly to the SCILAWS-BENCH work from yesterday, which flagged that existing evaluations can't distinguish genuine discovery from memorization. Here we see the same problem in a different domain: tabular models appear competent on physics datasets because those datasets contain noise and variation that reward pattern matching. The moment you remove the noise and ask for deterministic prediction, the illusion collapses. Both papers expose how benchmark design can hide fundamental reasoning gaps.
If TabPFN-3 or TabICLv2 are retrained with explicit unit tokens or symbolic constraint layers and recover performance on the deterministic subset, that confirms the gap is architectural rather than fundamental. If they remain stuck below 50% accuracy on noiseless equations even with these modifications, it signals these models may need hybrid symbolic-neural designs rather than pure scaling to handle physics applications.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsTabPFN-3 · TabICLv2 · TabDPT · Real-TabPFN-2.5
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as “Do Tabular Foundation Models Know Physics? Contamination, Units, and the Deterministic Limit”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.