Stress test reveals LLM performance collapse outside narrow optimization paths
Researchers have developed Decoding-Level Taboo, a runtime diagnostic that exposes gaps between LLM benchmark performance and real-world robustness by forcing models into unfamiliar generation paths. The method masks high-probability tokens during inference, compelling models to generate circumlocutions and revealing how quickly performance degrades when models deviate from their narrow optimization corridors. This stress test addresses a critical blind spot in model evaluation: most benchmarks measure performance under ideal conditions, not under the structural constraints and safety interventions that characterize production deployments. Early results across open-weight families suggest significant fragility lurking beneath published capability scores, with implications for deployment safety and the reliability gap between lab metrics and field performance.62











