
Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders
Sparse autoencoders have become central to mechanistic interpretability work, but their reliability hinges on a largely unexamined problem: feature reproducibility across training runs. This paper quantifies that instability through a per-feature stability metric, revealing a critical divide. Stable features encode genuine predictive signal and drive reconstruction quality, while unstable features are noise artifacts triggered by low-frequency surface patterns. The finding reshapes how researchers should weight SAE discoveries, suggesting many published interpretability claims rest on spurious, non-reproducible features. For labs building interpretability pipelines or relying on SAE-derived insights for alignment work, this establishes a filtering mechanism to separate robust mechanistic findings from statistical flukes.62

























