Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Isabelle, Lum, Joshua, Liu, Ziyi, Yogatama, Dani
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916457247408128
author Lee, Isabelle
Lum, Joshua
Liu, Ziyi
Yogatama, Dani
author_facet Lee, Isabelle
Lum, Joshua
Liu, Ziyi
Yogatama, Dani
contents While interpretability research has shed light on some internal algorithms utilized by transformer-based LLMs, reasoning in natural language, with its deep contextuality and ambiguity, defies easy categorization. As a result, formulating clear and motivating questions for circuit analysis that rely on well-defined in-domain and out-of-domain examples required for causal interventions is challenging. Although significant work has investigated circuits for specific tasks, such as indirect object identification (IOI), deciphering natural language reasoning through circuits remains difficult due to its inherent complexity. In this work, we take initial steps to characterize causal reasoning in LLMs by analyzing clear-cut cause-and-effect sentences like "I opened an umbrella because it started raining," where causal interventions may be possible through carefully crafted scenarios using GPT-2 small. Our findings indicate that causal syntax is localized within the first 2-3 layers, while certain heads in later layers exhibit heightened sensitivity to nonsensical variations of causal sentences. This suggests that models may infer reasoning by (1) detecting syntactic cues and (2) isolating distinct heads in the final layers that focus on semantic relationships.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21353
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics
Lee, Isabelle
Lum, Joshua
Liu, Ziyi
Yogatama, Dani
Computation and Language
Artificial Intelligence
While interpretability research has shed light on some internal algorithms utilized by transformer-based LLMs, reasoning in natural language, with its deep contextuality and ambiguity, defies easy categorization. As a result, formulating clear and motivating questions for circuit analysis that rely on well-defined in-domain and out-of-domain examples required for causal interventions is challenging. Although significant work has investigated circuits for specific tasks, such as indirect object identification (IOI), deciphering natural language reasoning through circuits remains difficult due to its inherent complexity. In this work, we take initial steps to characterize causal reasoning in LLMs by analyzing clear-cut cause-and-effect sentences like "I opened an umbrella because it started raining," where causal interventions may be possible through carefully crafted scenarios using GPT-2 small. Our findings indicate that causal syntax is localized within the first 2-3 layers, while certain heads in later layers exhibit heightened sensitivity to nonsensical variations of causal sentences. This suggests that models may infer reasoning by (1) detecting syntactic cues and (2) isolating distinct heads in the final layers that focus on semantic relationships.
title Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.21353