Mechanistic?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saphra, Naomi, Wiegreffe, Sarah
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913542919159808
author Saphra, Naomi
Wiegreffe, Sarah
author_facet Saphra, Naomi
Wiegreffe, Sarah
contents The rise of the term "mechanistic interpretability" has accompanied increasing interest in understanding neural models -- particularly language models. However, this jargon has also led to a fair amount of confusion. So, what does it mean to be "mechanistic"? We describe four uses of the term in interpretability research. The most narrow technical definition requires a claim of causality, while a broader technical definition allows for any exploration of a model's internals. However, the term also has a narrow cultural definition describing a cultural movement. To understand this semantic drift, we present a history of the NLP interpretability community and the formation of the separate, parallel "mechanistic" interpretability community. Finally, we discuss the broad cultural definition -- encompassing the entire field of interpretability -- and why the traditional NLP interpretability community has come to embrace it. We argue that the polysemy of "mechanistic" is the product of a critical divide within the interpretability community.
format Preprint
id arxiv_https___arxiv_org_abs_2410_09087
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mechanistic?
Saphra, Naomi
Wiegreffe, Sarah
Artificial Intelligence
Computation and Language
Machine Learning
The rise of the term "mechanistic interpretability" has accompanied increasing interest in understanding neural models -- particularly language models. However, this jargon has also led to a fair amount of confusion. So, what does it mean to be "mechanistic"? We describe four uses of the term in interpretability research. The most narrow technical definition requires a claim of causality, while a broader technical definition allows for any exploration of a model's internals. However, the term also has a narrow cultural definition describing a cultural movement. To understand this semantic drift, we present a history of the NLP interpretability community and the formation of the separate, parallel "mechanistic" interpretability community. Finally, we discuss the broad cultural definition -- encompassing the entire field of interpretability -- and why the traditional NLP interpretability community has come to embrace it. We argue that the polysemy of "mechanistic" is the product of a critical divide within the interpretability community.
title Mechanistic?
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.09087