Sparse confidence outputs corrupt LLM evaluation metrics, researchers warn
A new study exposes a critical flaw in how researchers evaluate LLM confidence estimates for classification tasks. The core problem: popular verbalization methods produce extremely sparse outputs, with models like Qwen3-32B collapsing confidence scores into just eight unique values across standard benchmarks. This sparsity doesn't merely limit real-world deployment; it corrupts evaluation itself. The researchers demonstrate that metric choice (stepwise versus linear interpolation in AUARC curves) can completely reverse method rankings, turning top performers into worst performers. The work calls for standardized evaluation protocols to prevent misleading comparisons and highlights a blind spot in how the field validates LLM reliability mechanisms.62

















