SAE evaluation results depend on which dictionary you use
A new paper exposes a hidden methodological flaw in how sparse autoencoders are evaluated. When researchers ablate a latent to measure its importance, they measure the effect at whichever token shows the strongest activation. This choice is typically made by the SAE itself, not the experimenter, and different dictionaries trained on the same model systematically pick different measurement points for similar latents. The finding undermines confidence in comparative SAE evaluations and suggests that published ablation results may reflect dictionary artifacts rather than true feature importance, forcing the interpretability community to reconsider how to benchmark these tools.
Modelwire context
ExplainerThe paper reveals that ablation position isn't a neutral experimental choice but a latent-specific artifact baked into each dictionary. Researchers don't consciously select where to measure; the SAE's training dynamics determine it, making published rankings of feature importance potentially incomparable across different sparse autoencoders trained on the same model.
This connects directly to the LittleLearner work from earlier this month, which tackled interpretability through controlled knowledge exposure. Both papers attack the same core problem: mechanistic interpretability research suffers from opaque, hard-to-reproduce experimental conditions. Where LittleLearner constrains the training signal to make knowledge acquisition observable, this paper exposes how evaluation methodology itself can hide systematic bias. Together they suggest the interpretability community needs to audit not just what models learn, but how we measure what they've learned.
If Google or Anthropic releases a revised SAE benchmark using position-agnostic ablation metrics (e.g., averaging effects across multiple token positions or using a held-out test set to select measurement points) within the next six months, that signals the community is taking this critique seriously. If published SAE rankings remain unchanged and new papers cite this work but don't change their evaluation protocol, the finding will have been noted but not acted upon.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsGoogle · sparse autoencoders · language models
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as “Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.