Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hanna, Michael, Pezzelle, Sandro, Belinkov, Yonatan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911955392921600
author Hanna, Michael
Pezzelle, Sandro
Belinkov, Yonatan
author_facet Hanna, Michael
Pezzelle, Sandro
Belinkov, Yonatan
contents Many recent language model (LM) interpretability studies have adopted the circuits framework, which aims to find the minimal computational subgraph, or circuit, that explains LM behavior on a given task. Most studies determine which edges belong in a LM's circuit by performing causal interventions on each edge independently, but this scales poorly with model size. Edge attribution patching (EAP), gradient-based approximation to interventions, has emerged as a scalable but imperfect solution to this problem. In this paper, we introduce a new method - EAP with integrated gradients (EAP-IG) - that aims to better maintain a core property of circuits: faithfulness. A circuit is faithful if all model edges outside the circuit can be ablated without changing the model's performance on the task; faithfulness is what justifies studying circuits, rather than the full model. Our experiments demonstrate that circuits found using EAP are less faithful than those found using EAP-IG, even though both have high node overlap with circuits found previously using causal interventions. We conclude more generally that when using circuits to compare the mechanisms models use to solve tasks, faithfulness, not overlap, is what should be measured.
format Preprint
id arxiv_https___arxiv_org_abs_2403_17806
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
Hanna, Michael
Pezzelle, Sandro
Belinkov, Yonatan
Machine Learning
Computation and Language
I.2.7
Many recent language model (LM) interpretability studies have adopted the circuits framework, which aims to find the minimal computational subgraph, or circuit, that explains LM behavior on a given task. Most studies determine which edges belong in a LM's circuit by performing causal interventions on each edge independently, but this scales poorly with model size. Edge attribution patching (EAP), gradient-based approximation to interventions, has emerged as a scalable but imperfect solution to this problem. In this paper, we introduce a new method - EAP with integrated gradients (EAP-IG) - that aims to better maintain a core property of circuits: faithfulness. A circuit is faithful if all model edges outside the circuit can be ablated without changing the model's performance on the task; faithfulness is what justifies studying circuits, rather than the full model. Our experiments demonstrate that circuits found using EAP are less faithful than those found using EAP-IG, even though both have high node overlap with circuits found previously using causal interventions. We conclude more generally that when using circuits to compare the mechanisms models use to solve tasks, faithfulness, not overlap, is what should be measured.
title Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
topic Machine Learning
Computation and Language
I.2.7
url https://arxiv.org/abs/2403.17806