A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Zhengxuan, Geiger, Atticus, Huang, Jing, Arora, Aryaman, Icard, Thomas, Potts, Christopher, Goodman, Noah D. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
by: Wu, Zhengxuan, et al.
Published: (2023)
by: Wu, Zhengxuan, et al.
Published: (2023)
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
by: Geiger, Atticus, et al.
Published: (2023)
by: Geiger, Atticus, et al.
Published: (2023)
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
by: Geiger, Atticus, et al.
Published: (2023)
by: Geiger, Atticus, et al.
Published: (2023)
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
ReFT: Representation Finetuning for Language Models
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
by: Huang, Jing, et al.
Published: (2024)
by: Huang, Jing, et al.
Published: (2024)
Bayesian scaling laws for in-context learning
by: Arora, Aryaman, et al.
Published: (2024)
by: Arora, Aryaman, et al.
Published: (2024)
How Causal Abstraction Underpins Computational Explanation
by: Geiger, Atticus, et al.
Published: (2025)
by: Geiger, Atticus, et al.
Published: (2025)
HyperSteer: Activation Steering at Scale with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
by: Puyin, Li, et al.
Published: (2026)
by: Puyin, Li, et al.
Published: (2026)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
Improved Representation Steering for Language Models
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
by: Arora, Aryaman, et al.
Published: (2024)
by: Arora, Aryaman, et al.
Published: (2024)
Language Model Circuits Are Sparse in the Neuron Basis
by: Arora, Aryaman, et al.
Published: (2026)
by: Arora, Aryaman, et al.
Published: (2026)
PreFT: Prefill-only finetuning for efficient inference
by: Lanpouthakoun, Andrew, et al.
Published: (2026)
by: Lanpouthakoun, Andrew, et al.
Published: (2026)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
Updating CLIP to Prefer Descriptions Over Captions
by: Zur, Amir, et al.
Published: (2024)
by: Zur, Amir, et al.
Published: (2024)
ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
by: Wu, Zhengxuan, et al.
Published: (2023)
by: Wu, Zhengxuan, et al.
Published: (2023)
How Do Transformers Learn Variable Binding in Symbolic Programs?
by: Wu, Yiwei, et al.
Published: (2025)
by: Wu, Yiwei, et al.
Published: (2025)
ADAG: Automatically Describing Attribution Graphs
by: Arora, Aryaman, et al.
Published: (2026)
by: Arora, Aryaman, et al.
Published: (2026)
Constructing Interpretable Features from Compositional Neuron Groups
by: Shafran, Or, et al.
Published: (2025)
by: Shafran, Or, et al.
Published: (2025)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
Mechanistic evaluation of Transformers and state space models
by: Arora, Aryaman, et al.
Published: (2025)
by: Arora, Aryaman, et al.
Published: (2025)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
by: She, Jingyuan Selena, et al.
Published: (2023)
by: She, Jingyuan Selena, et al.
Published: (2023)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
by: Zur, Amir, et al.
Published: (2025)
by: Zur, Amir, et al.
Published: (2025)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
by: Mishra, Shubhra, et al.
Published: (2024)
by: Mishra, Shubhra, et al.
Published: (2024)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
by: Zhong, Zexuan, et al.
Published: (2023)
by: Zhong, Zexuan, et al.
Published: (2023)
The Leaderboard Illusion
by: Singh, Shivalika, et al.
Published: (2025)
by: Singh, Shivalika, et al.
Published: (2025)
Evaluating and Optimizing Educational Content with Large Language Model Judgments
by: He-Yueya, Joy, et al.
Published: (2024)
by: He-Yueya, Joy, et al.
Published: (2024)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
by: Chaudhary, Maheep, et al.
Published: (2024)
by: Chaudhary, Maheep, et al.
Published: (2024)
Large Language Model Reasoning Failures
by: Song, Peiyang, et al.
Published: (2026)
by: Song, Peiyang, et al.
Published: (2026)
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
STaR-GATE: Teaching Language Models to Ask Clarifying Questions
by: Andukuri, Chinmaya, et al.
Published: (2024)
by: Andukuri, Chinmaya, et al.
Published: (2024)
Aligning to Illusions: Choice Blindness in Human and AI Feedback
by: Wu, Wenbin
Published: (2026)
by: Wu, Wenbin
Published: (2026)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
by: Bigelow, Eric, et al.
Published: (2026)
by: Bigelow, Eric, et al.
Published: (2026)
Is Child-Directed Speech Effective Training Data for Language Models?
by: Feng, Steven Y., et al.
Published: (2024)
by: Feng, Steven Y., et al.
Published: (2024)
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
by: Hase, Peter, et al.
Published: (2026)
by: Hase, Peter, et al.
Published: (2026)
Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
by: Boppana, Siddharth, et al.
Published: (2026)
by: Boppana, Siddharth, et al.
Published: (2026)
Similar Items
-
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
by: Wu, Zhengxuan, et al.
Published: (2023) -
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
by: Geiger, Atticus, et al.
Published: (2023) -
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
by: Geiger, Atticus, et al.
Published: (2023) -
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
by: Wu, Zhengxuan, et al.
Published: (2024) -
ReFT: Representation Finetuning for Language Models
by: Wu, Zhengxuan, et al.
Published: (2024)