Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nainani, Jatin, Vaidyanathan, Sankaran, Yeung, AJ, Gupta, Kartik, Jensen, David
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915049134620672
author Nainani, Jatin
Vaidyanathan, Sankaran
Yeung, AJ
Gupta, Kartik
Jensen, David
author_facet Nainani, Jatin
Vaidyanathan, Sankaran
Yeung, AJ
Gupta, Kartik
Jensen, David
contents Mechanistic interpretability aims to understand the inner workings of large neural networks by identifying circuits, or minimal subgraphs within the model that implement algorithms responsible for performing specific tasks. These circuits are typically discovered and analyzed using a narrowly defined prompt format. However, given the abilities of large language models (LLMs) to generalize across various prompt formats for the same task, it remains unclear how well these circuits generalize. For instance, it is unclear whether the models generalization results from reusing the same circuit components, the components behaving differently, or the use of entirely different components. In this paper, we investigate the generality of the indirect object identification (IOI) circuit in GPT-2 small, which is well-studied and believed to implement a simple, interpretable algorithm. We evaluate its performance on prompt variants that challenge the assumptions of this algorithm. Our findings reveal that the circuit generalizes surprisingly well, reusing all of its components and mechanisms while only adding additional input edges. Notably, the circuit generalizes even to prompt variants where the original algorithm should fail; we discover a mechanism that explains this which we term S2 Hacking. Our findings indicate that circuits within LLMs may be more flexible and general than previously recognized, underscoring the importance of studying circuit generalization to better understand the broader capabilities of these models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16105
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
Nainani, Jatin
Vaidyanathan, Sankaran
Yeung, AJ
Gupta, Kartik
Jensen, David
Machine Learning
Artificial Intelligence
Computation and Language
I.2.7
Mechanistic interpretability aims to understand the inner workings of large neural networks by identifying circuits, or minimal subgraphs within the model that implement algorithms responsible for performing specific tasks. These circuits are typically discovered and analyzed using a narrowly defined prompt format. However, given the abilities of large language models (LLMs) to generalize across various prompt formats for the same task, it remains unclear how well these circuits generalize. For instance, it is unclear whether the models generalization results from reusing the same circuit components, the components behaving differently, or the use of entirely different components. In this paper, we investigate the generality of the indirect object identification (IOI) circuit in GPT-2 small, which is well-studied and believed to implement a simple, interpretable algorithm. We evaluate its performance on prompt variants that challenge the assumptions of this algorithm. Our findings reveal that the circuit generalizes surprisingly well, reusing all of its components and mechanisms while only adding additional input edges. Notably, the circuit generalizes even to prompt variants where the original algorithm should fail; we discover a mechanism that explains this which we term S2 Hacking. Our findings indicate that circuits within LLMs may be more flexible and general than previously recognized, underscoring the importance of studying circuit generalization to better understand the broader capabilities of these models.
title Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
topic Machine Learning
Artificial Intelligence
Computation and Language
I.2.7
url https://arxiv.org/abs/2411.16105