A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Zhengxuan, Geiger, Atticus, Huang, Jing, Arora, Aryaman, Icard, Thomas, Potts, Christopher, Goodman, Noah D.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914649189908480
author Wu, Zhengxuan
Geiger, Atticus
Huang, Jing
Arora, Aryaman
Icard, Thomas
Potts, Christopher
Goodman, Noah D.
author_facet Wu, Zhengxuan
Geiger, Atticus
Huang, Jing
Arora, Aryaman
Icard, Thomas
Potts, Christopher
Goodman, Noah D.
contents We respond to the recent paper by Makelov et al. (2023), which reviews subspace interchange intervention methods like distributed alignment search (DAS; Geiger et al. 2023) and claims that these methods potentially cause "interpretability illusions". We first review Makelov et al. (2023)'s technical notion of what an "interpretability illusion" is, and then we show that even intuitive and desirable explanations can qualify as illusions in this sense. As a result, their method of discovering "illusions" can reject explanations they consider "non-illusory". We then argue that the illusions Makelov et al. (2023) see in practice are artifacts of their training and evaluation paradigms. We close by emphasizing that, though we disagree with their core characterization, Makelov et al. (2023)'s examples and discussion have undoubtedly pushed the field of interpretability forward.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12631
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
Wu, Zhengxuan
Geiger, Atticus
Huang, Jing
Arora, Aryaman
Icard, Thomas
Potts, Christopher
Goodman, Noah D.
Machine Learning
Artificial Intelligence
Computation and Language
We respond to the recent paper by Makelov et al. (2023), which reviews subspace interchange intervention methods like distributed alignment search (DAS; Geiger et al. 2023) and claims that these methods potentially cause "interpretability illusions". We first review Makelov et al. (2023)'s technical notion of what an "interpretability illusion" is, and then we show that even intuitive and desirable explanations can qualify as illusions in this sense. As a result, their method of discovering "illusions" can reject explanations they consider "non-illusory". We then argue that the illusions Makelov et al. (2023) see in practice are artifacts of their training and evaluation paradigms. We close by emphasizing that, though we disagree with their core characterization, Makelov et al. (2023)'s examples and discussion have undoubtedly pushed the field of interpretability forward.
title A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2401.12631