Pruning breaks sparse autoencoders unless activation geometry is preserved
Researchers have identified a critical vulnerability in how model pruning interacts with sparse autoencoders, the primary tool for mechanistic interpretability of LLMs. The work shows that standard magnitude-based pruning corrupts the learned representation geometry that SAEs depend on, while activation-aware methods like Wanda and SparseGPT implicitly preserve this structure through perturbation energy control. This finding matters because it exposes a gap between two major LLM research directions: as practitioners compress models for deployment, interpretability tools become unreliable, potentially undermining safety audits and mechanistic understanding efforts that rely on SAE analysis.62









