Modelwire
Subscribe

Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders

Illustration accompanying: Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders

Sparse autoencoders have become central to mechanistic interpretability work, but their reliability hinges on a largely unexamined problem: feature reproducibility across training runs. This paper quantifies that instability through a per-feature stability metric, revealing a critical divide. Stable features encode genuine predictive signal and drive reconstruction quality, while unstable features are noise artifacts triggered by low-frequency surface patterns. The finding reshapes how researchers should weight SAE discoveries, suggesting many published interpretability claims rest on spurious, non-reproducible features. For labs building interpretability pipelines or relying on SAE-derived insights for alignment work, this establishes a filtering mechanism to separate robust mechanistic findings from statistical flukes.

Modelwire context

Explainer

The deeper provocation here isn't just that some SAE features are noisy: it's that the field has been publishing mechanistic claims without a standard for checking whether those claims survive a different random seed. This paper is essentially proposing the first quality filter for a body of literature that has been accumulating without one.

The related coverage from this period is largely disconnected from this story. The nD-RoPE work and the sea surface temperature forecasting paper both touch on representational geometry and dimensionality, but neither engages with interpretability pipelines or SAE training dynamics. The more relevant context sits outside the current archive: the broader mechanistic interpretability push at labs like Anthropic, which has treated SAEs as a primary lens for understanding residual stream features. That body of work is what this paper implicitly audits.

Watch whether Anthropic or any lab with a public SAE release applies this stability metric retroactively to their published feature dictionaries and reports what fraction of claimed features survive the filter. If that number comes in below 50 percent, it will force a significant reassessment of which prior interpretability findings are worth building on.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsSparse Autoencoders · Neural Network Interpretability

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders · Modelwire