
LLM benchmark stability masks per-example prediction instability
A new study reveals that state-of-the-art LLMs mask fragility behind stable aggregate benchmarks. While overall accuracy remains unchanged when task-irrelevant context is prepended to questions, individual predictions flip unpredictably on a subset of examples, even when triggered by meaningless character sequences. This instability persists across multiple model families and datasets, suggesting that current evaluation metrics fail to capture real-world brittleness in context-rich deployments. The finding challenges assumptions about model robustness and has direct implications for production systems relying on benchmark scores as reliability proxies.62





















