Demystifying Variance in Circuit Discovery of LLMs

Mechanistic interpretability research has long struggled with a fundamental problem: circuit discovery methods that identify which model components drive specific behaviors produce inconsistent results across different data batches, prompt phrasings, and individual samples. This paper diagnoses the root causes of this variance in the current state-of-the-art EAP-IG method and proposes CEAP as an improved alternative. The work matters because unreliable circuit discovery undermines the entire interpretability agenda, making it harder for researchers to build trustworthy explanations of LLM reasoning. Solving variance is a prerequisite for mechanistic interpretability to move from academic exercise to practical tool for safety and debugging.
Modelwire context
ExplainerThe paper's contribution isn't a faster circuit discovery method but a diagnostic one: it identifies *why* existing methods produce different answers on different runs, which is a prerequisite problem that has quietly invalidated a lot of prior interpretability work without anyone formally accounting for it.
This pairs directly with 'Scalable Circuit Learning for Interpreting Large Language Models' from the same day, which introduced CircuitLasso as a compute-efficient alternative to intervention-based circuit discovery. That paper solved the cost problem; this one addresses the reliability problem. Together they sketch a two-front effort to make circuit discovery actually usable: cheaper and more consistent. Neither paper alone closes the loop, but the convergence of both on the same date suggests the interpretability community is actively stress-testing the foundations of the method rather than just applying it to new models.
Watch whether CEAP gets adopted as a baseline in follow-on circuit discovery work over the next six months. If CircuitLasso or similar scalable methods begin citing CEAP for variance control, that signals the field is converging on a shared methodology rather than continuing to produce incomparable results.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.