Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Geiger, Atticus, Ibeling, Duligur, Zur, Amir, Chaudhary, Maheep, Chauhan, Sonakshi, Huang, Jing, Arora, Aryaman, Wu, Zhengxuan, Goodman, Noah, Potts, Christopher, Icard, Thomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
von: Geiger, Atticus, et al.
Veröffentlicht: (2023)
von: Geiger, Atticus, et al.
Veröffentlicht: (2023)
On Probabilistic and Causal Reasoning with Summation Operators
von: Ibeling, Duligur, et al.
Veröffentlicht: (2024)
von: Ibeling, Duligur, et al.
Veröffentlicht: (2024)
How Causal Abstraction Underpins Computational Explanation
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2024)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2024)
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
von: Huang, Jing, et al.
Veröffentlicht: (2024)
von: Huang, Jing, et al.
Veröffentlicht: (2024)
ReFT: Representation Finetuning for Language Models
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
von: Puyin, Li, et al.
Veröffentlicht: (2026)
von: Puyin, Li, et al.
Veröffentlicht: (2026)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
Activation Steering via Generative Causal Mediation
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)
von: Zur, Amir, et al.
Veröffentlicht: (2024)
Bayesian scaling laws for in-context learning
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
Punctuation and Predicates in Language Models
von: Chauhan, Sonakshi, et al.
Veröffentlicht: (2025)
von: Chauhan, Sonakshi, et al.
Veröffentlicht: (2025)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
Combining Causal Models for More Accurate Abstractions of Neural Networks
von: Pîslar, Theodora-Mara, et al.
Veröffentlicht: (2025)
von: Pîslar, Theodora-Mara, et al.
Veröffentlicht: (2025)
Improved Representation Steering for Language Models
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
von: Zur, Amir, et al.
Veröffentlicht: (2025)
von: Zur, Amir, et al.
Veröffentlicht: (2025)
HyperSteer: Activation Steering at Scale with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
Modeling Discrimination with Causal Abstraction
von: Mossé, Milan, et al.
Veröffentlicht: (2025)
von: Mossé, Milan, et al.
Veröffentlicht: (2025)
In-Context Environments Induce Evaluation-Awareness in Language Models
von: Chaudhary, Maheep
Veröffentlicht: (2026)
von: Chaudhary, Maheep
Veröffentlicht: (2026)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
Mechanistic evaluation of Transformers and state space models
von: Arora, Aryaman, et al.
Veröffentlicht: (2025)
von: Arora, Aryaman, et al.
Veröffentlicht: (2025)
PreFT: Prefill-only finetuning for efficient inference
von: Lanpouthakoun, Andrew, et al.
Veröffentlicht: (2026)
von: Lanpouthakoun, Andrew, et al.
Veröffentlicht: (2026)
Constructing Interpretable Features from Compositional Neuron Groups
von: Shafran, Or, et al.
Veröffentlicht: (2025)
von: Shafran, Or, et al.
Veröffentlicht: (2025)
ADAG: Automatically Describing Attribution Graphs
von: Arora, Aryaman, et al.
Veröffentlicht: (2026)
von: Arora, Aryaman, et al.
Veröffentlicht: (2026)
Language Model Circuits Are Sparse in the Neuron Basis
von: Arora, Aryaman, et al.
Veröffentlicht: (2026)
von: Arora, Aryaman, et al.
Veröffentlicht: (2026)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
von: Huang, Jing, et al.
Veröffentlicht: (2025)
von: Huang, Jing, et al.
Veröffentlicht: (2025)
ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
MIB: A Mechanistic Interpretability Benchmark
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
von: Sutter, Denis, et al.
Veröffentlicht: (2025)
von: Sutter, Denis, et al.
Veröffentlicht: (2025)
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
von: Chaudhary, Siddharth, et al.
Veröffentlicht: (2025)
von: Chaudhary, Siddharth, et al.
Veröffentlicht: (2025)
Transcoder Adapters for Reasoning-Model Diffing
von: Hu, Nathan, et al.
Veröffentlicht: (2026)
von: Hu, Nathan, et al.
Veröffentlicht: (2026)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
How Do Transformers Learn Variable Binding in Symbolic Programs?
von: Wu, Yiwei, et al.
Veröffentlicht: (2025)
von: Wu, Yiwei, et al.
Veröffentlicht: (2025)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
von: Nunez, Jeanmely Rojas, et al.
Veröffentlicht: (2026)
von: Nunez, Jeanmely Rojas, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023) -
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024) -
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
von: Geiger, Atticus, et al.
Veröffentlicht: (2023) -
On Probabilistic and Causal Reasoning with Summation Operators
von: Ibeling, Duligur, et al.
Veröffentlicht: (2024) -
How Causal Abstraction Underpins Computational Explanation
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)