
On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
Researchers have identified a root cause of feature death in sparse autoencoders, a critical interpretability tool for decomposing neural network behavior. The problem, where learned features never activate and waste dictionary capacity, stems from activation outliers that permanently suppress certain features at initialization. By formalizing outlier severity as a ratio of mean to variance magnitude, the work explains why death rates swing from near-zero in GPT-2 to over 70% in AlphaFold3 under identical configurations. This finding matters for mechanistic interpretability efforts and SAE reliability across diverse model architectures.62


























