A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914649189908480 |
|---|---|
| author | Wu, Zhengxuan Geiger, Atticus Huang, Jing Arora, Aryaman Icard, Thomas Potts, Christopher Goodman, Noah D. |
| author_facet | Wu, Zhengxuan Geiger, Atticus Huang, Jing Arora, Aryaman Icard, Thomas Potts, Christopher Goodman, Noah D. |
| contents | We respond to the recent paper by Makelov et al. (2023), which reviews subspace interchange intervention methods like distributed alignment search (DAS; Geiger et al. 2023) and claims that these methods potentially cause "interpretability illusions". We first review Makelov et al. (2023)'s technical notion of what an "interpretability illusion" is, and then we show that even intuitive and desirable explanations can qualify as illusions in this sense. As a result, their method of discovering "illusions" can reject explanations they consider "non-illusory". We then argue that the illusions Makelov et al. (2023) see in practice are artifacts of their training and evaluation paradigms. We close by emphasizing that, though we disagree with their core characterization, Makelov et al. (2023)'s examples and discussion have undoubtedly pushed the field of interpretability forward. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_12631 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments Wu, Zhengxuan Geiger, Atticus Huang, Jing Arora, Aryaman Icard, Thomas Potts, Christopher Goodman, Noah D. Machine Learning Artificial Intelligence Computation and Language We respond to the recent paper by Makelov et al. (2023), which reviews subspace interchange intervention methods like distributed alignment search (DAS; Geiger et al. 2023) and claims that these methods potentially cause "interpretability illusions". We first review Makelov et al. (2023)'s technical notion of what an "interpretability illusion" is, and then we show that even intuitive and desirable explanations can qualify as illusions in this sense. As a result, their method of discovering "illusions" can reject explanations they consider "non-illusory". We then argue that the illusions Makelov et al. (2023) see in practice are artifacts of their training and evaluation paradigms. We close by emphasizing that, though we disagree with their core characterization, Makelov et al. (2023)'s examples and discussion have undoubtedly pushed the field of interpretability forward. |
| title | A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2401.12631 |