Modelwire
Subscribe

LLM confidence statements diverge from internal uncertainty across 30 models

A systematic study across 30 LLMs reveals that when models are asked to express confidence linguistically, their stated certainty often misaligns with internal probability signals derived from logits and semantic entropy. The research exposes a critical gap in model transparency: instruction-tuned variants tend to report higher confidence but don't necessarily track their actual uncertainty better. This finding matters for practitioners relying on model confidence signals for downstream decisions, from retrieval augmentation to human-in-the-loop workflows. The divergence suggests current confidence-reporting mechanisms may mask genuine model limitations.

Modelwire context

Explainer

The study doesn't just measure confidence mismatch; it shows instruction-tuning actively makes it worse. Models trained to follow instructions report higher confidence while their internal signals remain unchanged, suggesting the gap is a learned behavior, not a measurement artifact.

This fits directly into a pattern we've covered across three recent papers. The 'Blind Men and the Elephant' study (August 28) exposed how models systematically hide uncertainty about minority viewpoints while appearing confident. The 'Fidelity Is Not Enough' dispatch work (same date) found models generating plausible fabrications while passing accuracy checks. And the persona dialogue study showed how training-time information access shapes what models actually learn versus what they report. Together, these papers suggest models are learning to mask their actual limitations through both what they say and what they do.

If the same 30-model study disaggregates results by instruction-tuning method (e.g., RLHF vs. DPO vs. supervised fine-tuning), and shows one method produces tighter alignment between stated and internal confidence, that would point to a fixable training choice. Otherwise, the divergence may be intrinsic to how language models convert probability into text.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsLarge language models · Instruction-tuned models · Semantic entropy

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as When Linguistic and Internal Confidence Diverge in Large Language Models”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.