Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Edin, Joakim, Christensen, Casper L., Csordás, Róbert, Ruotsalo, Tuukka, Wu, Zhengxuan, Maistro, Maria, Huang, Jing, Maaløe, Lars
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911740062597120
author Edin, Joakim
Christensen, Casper L.
Csordás, Róbert
Ruotsalo, Tuukka
Wu, Zhengxuan
Maistro, Maria
Huang, Jing
Maaløe, Lars
author_facet Edin, Joakim
Christensen, Casper L.
Csordás, Róbert
Ruotsalo, Tuukka
Wu, Zhengxuan
Maistro, Maria
Huang, Jing
Maaløe, Lars
contents Circuit localization methods aim to identify the subset of model components responsible for specific behaviors in large language models, enabling detailed mechanistic analysis. Most existing methods assume components act independently and estimate importance by perturbing each component in isolation. However, components in neural networks interact, and ignoring these interactions leads to systematic misestimation of component importance. We find that one particularly problematic interaction is attention self-repair, in which softmax redistribution causes gradients for influential attention scores to vanish as other positions with similar values compensate. We introduce Gradient Interaction Modifications (GIM), a technique that explicitly accounts for feature interactions during backpropagation. GIM achieves state-of-the-art performance on the circuit localization track of the Mechanistic Interpretability Benchmark and outperforms existing gradient-based methods on feature attribution across diverse tasks. By accounting for interaction effects and explaining why prior methods underestimate component importance, GIM enables more faithful mechanistic analysis of large language models. GIM is available as a Python package at https://github.com/corticph/gim.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17630
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
Edin, Joakim
Christensen, Casper L.
Csordás, Róbert
Ruotsalo, Tuukka
Wu, Zhengxuan
Maistro, Maria
Huang, Jing
Maaløe, Lars
Computation and Language
Machine Learning
68T07
I.2.0; I.2.7
Circuit localization methods aim to identify the subset of model components responsible for specific behaviors in large language models, enabling detailed mechanistic analysis. Most existing methods assume components act independently and estimate importance by perturbing each component in isolation. However, components in neural networks interact, and ignoring these interactions leads to systematic misestimation of component importance. We find that one particularly problematic interaction is attention self-repair, in which softmax redistribution causes gradients for influential attention scores to vanish as other positions with similar values compensate. We introduce Gradient Interaction Modifications (GIM), a technique that explicitly accounts for feature interactions during backpropagation. GIM achieves state-of-the-art performance on the circuit localization track of the Mechanistic Interpretability Benchmark and outperforms existing gradient-based methods on feature attribution across diverse tasks. By accounting for interaction effects and explaining why prior methods underestimate component importance, GIM enables more faithful mechanistic analysis of large language models. GIM is available as a Python package at https://github.com/corticph/gim.
title Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
topic Computation and Language
Machine Learning
68T07
I.2.0; I.2.7
url https://arxiv.org/abs/2505.17630