Modelwire
Subscribe

Time-series foundation models fail hold-out test, matching naive baselines

Researchers constructed a contamination-free evaluation framework for time-series foundation models by building a hold-out test set using only data published after model release dates. Testing six pretrained forecasters alongside classical and dataset-specific baselines across seven groups revealed a sobering finding: pretrained models won marginally but showed no advantage on daily exchange rates, matching naive seasonal forecasts. This work exposes a critical gap between benchmark performance and real-world utility, forcing the field to reckon with whether pretraining gains on public archives reflect genuine capability or merely test-set leakage. The negative result matters because it challenges assumptions baked into foundation model evaluation across domains.

Modelwire context

Skeptical read

The real finding isn't that pretraining sometimes fails on held-out data. It's that a contamination-free evaluation framework reveals pretraining may have been winning on benchmarks precisely because those benchmarks were contaminated, not because the models learned transferable forecasting skill.

This connects to the federated learning work on OmniMed-FL from the same day, which also grapples with evaluation rigor but from the opposite angle: that paper builds frameworks to enable training across distributed data without leakage. Both papers treat data isolation as a first-class design problem rather than an afterthought. The difference is OmniMed-FL solves for privacy and compliance, while this time-series work solves for benchmark validity. Together they suggest the field is finally taking seriously what happens when you actually enforce the boundary between training and test.

If Theta or other foundation model vendors release new time-series models trained explicitly on data before a fixed cutoff date, and those models are then evaluated only on post-cutoff data, watch whether their gains over classical baselines shrink to the margins reported here. If they don't, it signals the contamination problem is specific to existing archives rather than a systemic issue with pretraining for forecasting.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsTheta · Time-series foundation models

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as A Later Test Set Is Not a New Domain: Pretraining Familiarity Survives a Contamination-Free Hold-Out”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Time-series foundation models fail hold-out test, matching naive baselines · Modelwire