Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Lin, Hu, Lijie, Wang, Di
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909493021900800
author Zhang, Lin
Hu, Lijie
Wang, Di
author_facet Zhang, Lin
Hu, Lijie
Wang, Di
contents Transformer-based language models have achieved significant success; however, their internal mechanisms remain largely opaque due to the complexity of non-linear interactions and high-dimensional operations. While previous studies have demonstrated that these models implicitly embed reasoning trees, humans typically employ various distinct logical reasoning mechanisms to complete the same task. It is still unclear which multi-step reasoning mechanisms are used by language models to solve such tasks. In this paper, we aim to address this question by investigating the mechanistic interpretability of language models, particularly in the context of multi-step reasoning tasks. Specifically, we employ circuit analysis and self-influence functions to evaluate the changing importance of each token throughout the reasoning process, allowing us to map the reasoning paths adopted by the model. We apply this methodology to the GPT-2 model on a prediction task (IOI) and demonstrate that the underlying circuits reveal a human-interpretable reasoning process used by the model.
format Preprint
id arxiv_https___arxiv_org_abs_2502_09022
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning
Zhang, Lin
Hu, Lijie
Wang, Di
Artificial Intelligence
Transformer-based language models have achieved significant success; however, their internal mechanisms remain largely opaque due to the complexity of non-linear interactions and high-dimensional operations. While previous studies have demonstrated that these models implicitly embed reasoning trees, humans typically employ various distinct logical reasoning mechanisms to complete the same task. It is still unclear which multi-step reasoning mechanisms are used by language models to solve such tasks. In this paper, we aim to address this question by investigating the mechanistic interpretability of language models, particularly in the context of multi-step reasoning tasks. Specifically, we employ circuit analysis and self-influence functions to evaluate the changing importance of each token throughout the reasoning process, allowing us to map the reasoning paths adopted by the model. We apply this methodology to the GPT-2 model on a prediction task (IOI) and demonstrate that the underlying circuits reveal a human-interpretable reasoning process used by the model.
title Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning
topic Artificial Intelligence
url https://arxiv.org/abs/2502.09022