How Causal Abstraction Underpins Computational Explanation
Fuente:
arXiv
Saved in:
| Main Authors: | Geiger, Atticus, Harding, Jacqueline, Icard, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
by: Puyin, Li, et al.
Published: (2026)
by: Puyin, Li, et al.
Published: (2026)
How Do Transformers Learn Variable Binding in Symbolic Programs?
by: Wu, Yiwei, et al.
Published: (2025)
by: Wu, Yiwei, et al.
Published: (2025)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
Combining Causal Models for More Accurate Abstractions of Neural Networks
by: Pîslar, Theodora-Mara, et al.
Published: (2025)
by: Pîslar, Theodora-Mara, et al.
Published: (2025)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
HyperSteer: Activation Steering at Scale with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
What is it for a Machine Learning Model to Have a Capability?
by: Harding, Jacqueline, et al.
Published: (2024)
by: Harding, Jacqueline, et al.
Published: (2024)
Causal Abstraction in Model Interpretability: A Compact Survey
by: Zhang, Yihao
Published: (2024)
by: Zhang, Yihao
Published: (2024)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
by: Chaudhary, Maheep, et al.
Published: (2024)
by: Chaudhary, Maheep, et al.
Published: (2024)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
ReFT: Representation Finetuning for Language Models
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
by: Bigelow, Eric, et al.
Published: (2026)
by: Bigelow, Eric, et al.
Published: (2026)
Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
by: Boppana, Siddharth, et al.
Published: (2026)
by: Boppana, Siddharth, et al.
Published: (2026)
CausalARC: Abstract Reasoning with Causal World Models
by: Maasch, Jacqueline, et al.
Published: (2025)
by: Maasch, Jacqueline, et al.
Published: (2025)
From Words to Actions: Unveiling the Theoretical Underpinnings of LLM-Driven Autonomous Systems
by: He, Jianliang, et al.
Published: (2024)
by: He, Jianliang, et al.
Published: (2024)
Constructing Interpretable Features from Compositional Neuron Groups
by: Shafran, Or, et al.
Published: (2025)
by: Shafran, Or, et al.
Published: (2025)
Modeling Discrimination with Causal Abstraction
by: Mossé, Milan, et al.
Published: (2025)
by: Mossé, Milan, et al.
Published: (2025)
How Reliable are Causal Probing Interventions?
by: Canby, Marc, et al.
Published: (2024)
by: Canby, Marc, et al.
Published: (2024)
Capturing Sparks of Abstraction for the ARC Challenge
by: Andrews, Martin
Published: (2024)
by: Andrews, Martin
Published: (2024)
Show Me How It's Done: The Role of Explanations in Fine-Tuning Language Models
by: Ballout, Mohamad, et al.
Published: (2024)
by: Ballout, Mohamad, et al.
Published: (2024)
Multiple Abstraction Level Retrieve Augment Generation
by: Zheng, Zheng, et al.
Published: (2025)
by: Zheng, Zheng, et al.
Published: (2025)
Compositional Causal Reasoning Evaluation in Language Models
by: Maasch, Jacqueline R. M. A., et al.
Published: (2025)
by: Maasch, Jacqueline R. M. A., et al.
Published: (2025)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
by: Wu, Zhengxuan, et al.
Published: (2023)
by: Wu, Zhengxuan, et al.
Published: (2023)
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
by: Geiger, Atticus, et al.
Published: (2023)
by: Geiger, Atticus, et al.
Published: (2023)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
by: Qu, Yuxiao, et al.
Published: (2025)
by: Qu, Yuxiao, et al.
Published: (2025)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Parallel LLM Reasoning for Bias-Resilient, Robust Conceptual Abstraction
by: Adeseye, Aisvarya, et al.
Published: (2026)
by: Adeseye, Aisvarya, et al.
Published: (2026)
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
by: Geiger, Atticus, et al.
Published: (2023)
by: Geiger, Atticus, et al.
Published: (2023)
Activation Steering via Generative Causal Mediation
by: Sankaranarayanan, Aruna, et al.
Published: (2026)
by: Sankaranarayanan, Aruna, et al.
Published: (2026)
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
by: Khan, Zaid, et al.
Published: (2025)
by: Khan, Zaid, et al.
Published: (2025)
Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
by: Chen, Yanda, et al.
Published: (2024)
by: Chen, Yanda, et al.
Published: (2024)
Exposing propaganda: an analysis of stylistic cues comparing human annotations and machine classification
by: Faye, Géraud, et al.
Published: (2024)
by: Faye, Géraud, et al.
Published: (2024)
STELLA: Guiding Large Language Models for Time Series Forecasting with Semantic Abstractions
by: Fan, Junjie, et al.
Published: (2025)
by: Fan, Junjie, et al.
Published: (2025)
How do LLMs Compute Verbal Confidence
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
Causal Estimation of Tokenisation Bias
by: Lesci, Pietro, et al.
Published: (2025)
by: Lesci, Pietro, et al.
Published: (2025)
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models
by: Zheng, Huaixiu Steven, et al.
Published: (2023)
by: Zheng, Huaixiu Steven, et al.
Published: (2023)
When Should LLMs Be Less Specific? Selective Abstraction for Reliable Long-Form Text Generation
by: Goren, Shani, et al.
Published: (2026)
by: Goren, Shani, et al.
Published: (2026)
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
by: Damani, Mehul, et al.
Published: (2024)
by: Damani, Mehul, et al.
Published: (2024)
Similar Items
-
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
by: Puyin, Li, et al.
Published: (2026) -
How Do Transformers Learn Variable Binding in Symbolic Programs?
by: Wu, Yiwei, et al.
Published: (2025) -
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
by: Wu, Zhengxuan, et al.
Published: (2024) -
Combining Causal Models for More Accurate Abstractions of Neural Networks
by: Pîslar, Theodora-Mara, et al.
Published: (2025) -
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
by: Huang, Jing, et al.
Published: (2025)