Modelwire
Subscribe

Cough-based TB models fail to generalize across recording devices

A systematic evaluation of machine learning models trained on cough acoustics for tuberculosis screening reveals a critical generalization failure across datasets. While individual models achieved moderate within-dataset performance (ROC-AUC 0.755), external validation dropped below 0.6, suggesting models learned device and collection artifacts rather than disease signals. The finding exposes a fundamental challenge in healthcare ML: acoustic representations clustered by recording equipment and geography rather than clinical phenotype, with device mismatch severely degrading transfer learning. This work highlights why real-world deployment of audio-based diagnostic systems requires rethinking data collection and model architecture, not just larger datasets.

Modelwire context

Explainer

The paper's core insight isn't just that models fail on new data (known problem), but that they fail because acoustic representations cluster by recording device and geography rather than by disease phenotype. This means the models never learned TB signals at all.

This joins a pattern across recent ML research: benchmarks and datasets are masking fundamental generalization failures. The Key Point Analysis work from late August exposed how flawed benchmarks let models appear capable while hiding real ceilings. Here, within-dataset ROC-AUC of 0.755 looked acceptable until external validation dropped below 0.6, revealing the benchmark itself was the artifact. The same dynamic appears in fact-checking systems evaluated across domains last month, where single-benchmark fine-tuning created an illusion of capability that evaporated on transfer. The difference: those papers proposed fixes (better benchmarks, domain-aware architectures). This TB screening work stops at diagnosis.

If the authors or follow-up work show that models trained on deliberately device-diverse data (mixing smartphone, clinical microphone, and field recordings from the start) maintain above 0.70 ROC-AUC on held-out device types, that confirms device mismatch was the root cause. If performance stays below 0.65 even with mixed-device training, the problem is deeper (TB cough signals may be too subtle or too confounded by non-disease factors to extract acoustically).

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsCODA · tuberculosis screening · machine learning · deep learning

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as Why ML-based cough models do not generalize: a systematic cross-dataset evaluation for tuberculosis screening”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Cough-based TB models fail to generalize across recording devices · Modelwire