
Super Weights fail to improve LLMs when trained in isolation
A new study challenges the premise that Super Weights, individual parameters claimed to be disproportionately important in LLMs, are actually critical to model function. Researchers found that pruning Super Weights does not consistently harm performance across different models, and counterintuitively, training these supposedly vital parameters in isolation causes catastrophic accuracy collapse. Training random parameters in the same layers instead maintains baseline performance, suggesting Super Weight identification may reflect statistical artifacts rather than genuine architectural bottlenecks. This finding undermines recent pruning and sparsity research that relied on Super Weight targeting, forcing a recalibration of how practitioners think about parameter importance and selective training strategies.62
























