Modelwire
Subscribe

Cardiovascular screening models inflated by target leakage, not foundation model gains

A systematic audit of ten machine learning classifiers reveals that reported 0.89 AUROC scores in cardiovascular screening models stem primarily from target leakage rather than genuine predictive power. The study benchmarks linear, tree-ensemble, neural, glass-box, and tabular foundation models across five feature tiers on 442,067 survey respondents, evaluating discrimination, calibration, fairness, conformal coverage, and explanation fidelity. This work exposes a critical gap between published accuracy claims and deployment-ready performance in healthcare AI, with implications for how foundation models are validated in regulated domains where leakage-driven inflation can mask poor real-world generalization.

Modelwire context

Explainer

The study's core finding isn't that leakage exists in ML (known for years), but that it explains nearly all reported accuracy gains in cardiovascular screening, and that this pattern persists across both classical and foundation model architectures. The implication: model novelty masks a validation problem, not a capability problem.

This audit echoes a pattern from two recent studies. The CGM forecasting work (September) showed foundation models don't automatically transfer to medical tasks without adaptation, revealing that general pretraining doesn't guarantee domain performance. The CausalArena benchmark (same period) flagged how fixed evaluation protocols conflate genuine reasoning with data memorization. This cardiovascular screening audit extends that concern: reported 0.89 AUROC scores across multiple model classes suggest the benchmark itself is contaminated, not that any single architecture solved the problem. The common thread is that inflated metrics hide deployment failures.

If the authors retrain these ten classifiers on a held-out temporal split (models trained on 2020-2022 data, tested on 2023-2024 BRFSS respondents), watch whether AUROC drops below 0.75. If it does, leakage is confirmed as the primary driver. If performance holds, the issue is more subtle (e.g., stable confounding rather than target leakage), and the field needs a different fix.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsBehavioral Risk Factor Surveillance System · tabular foundation models · glass-box models

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.