MedCLIP shortcuts undermine chest X-ray model robustness across layers
Researchers have exposed a critical vulnerability in medical vision-language models: CLIP-based systems like MedCLIP exploit dataset shortcuts rather than learning robust diagnostic patterns. By instrumenting ResNet-50 with 17 probes across layers, the team traced how models achieve high accuracy on chest X-ray tasks (pneumothorax, cardiomegaly) while relying on spurious correlations invisible to standard evaluation. This finding matters because medical AI deployment assumes learned features generalize across hospitals and populations. The layer-wise analysis reveals shortcuts emerge early and persist, suggesting current calibration methods mask brittle decision-making. For practitioners, this signals that SOTA metrics on benchmark datasets may not predict real-world reliability in clinical settings.62















