
The strength of clinical evidence is recoverable from language model representations but not from their stated grades
Researchers tested whether 22 open-weight LLMs can internally represent clinical evidence strength, separate from factual accuracy. Using 45,134 harmonized medical claims across three grading frameworks, they found that linear probes successfully recovered evidence grades from model activations in every tested model, despite the systems rarely stating confidence levels explicitly when queried. This gap between hidden representational capacity and stated outputs has direct implications for clinical AI deployment, where confidence calibration failures could propagate silently through downstream applications.62




























