The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sutter, Denis, Minder, Julian, Hofmann, Thomas, Pimentel, Tiago |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Causal Estimation of Memorisation Profiles
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
Causal Abstraction Inference under Lossy Representations
von: Xia, Kevin, et al.
Veröffentlicht: (2025)
von: Xia, Kevin, et al.
Veröffentlicht: (2025)
Learning Causal Abstractions of Linear Structural Causal Models
von: Massidda, Riccardo, et al.
Veröffentlicht: (2024)
von: Massidda, Riccardo, et al.
Veröffentlicht: (2024)
Causal Estimation of Tokenisation Bias
von: Lesci, Pietro, et al.
Veröffentlicht: (2025)
von: Lesci, Pietro, et al.
Veröffentlicht: (2025)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
von: Minder, Julian, et al.
Veröffentlicht: (2025)
von: Minder, Julian, et al.
Veröffentlicht: (2025)
Causal Abstraction in Model Interpretability: A Compact Survey
von: Zhang, Yihao
Veröffentlicht: (2024)
von: Zhang, Yihao
Veröffentlicht: (2024)
On the Identifiability of Causal Abstractions
von: Li, Xiusi, et al.
Veröffentlicht: (2025)
von: Li, Xiusi, et al.
Veröffentlicht: (2025)
Challenges in Mechanistically Interpreting Model Representations
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
Causal Abstractions, Categorically Unified
von: Englberger, Markus, et al.
Veröffentlicht: (2025)
von: Englberger, Markus, et al.
Veröffentlicht: (2025)
Neural Causal Abstractions
von: Xia, Kevin, et al.
Veröffentlicht: (2024)
von: Xia, Kevin, et al.
Veröffentlicht: (2024)
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
von: Roy, Dip, et al.
Veröffentlicht: (2025)
von: Roy, Dip, et al.
Veröffentlicht: (2025)
Distributionally Robust Causal Abstractions
von: Felekis, Yorgos, et al.
Veröffentlicht: (2025)
von: Felekis, Yorgos, et al.
Veröffentlicht: (2025)
How Causal Abstraction Underpins Computational Explanation
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
Open Problems in Mechanistic Interpretability
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
Exemplar Partitioning for Mechanistic Interpretability
von: Rumbelow, Jessica
Veröffentlicht: (2026)
von: Rumbelow, Jessica
Veröffentlicht: (2026)
In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention
von: He, Jianliang, et al.
Veröffentlicht: (2025)
von: He, Jianliang, et al.
Veröffentlicht: (2025)
The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
von: Mueller, Aaron, et al.
Veröffentlicht: (2024)
von: Mueller, Aaron, et al.
Veröffentlicht: (2024)
The Causal Information Bottleneck and Optimal Causal Variable Abstractions
von: Simoes, Francisco N. F. Q., et al.
Veröffentlicht: (2024)
von: Simoes, Francisco N. F. Q., et al.
Veröffentlicht: (2024)
On the Effect of (Near) Duplicate Subwords in Language Modelling
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
Learning Consistent Causal Abstraction Networks
von: D'Acunto, Gabriele, et al.
Veröffentlicht: (2026)
von: D'Acunto, Gabriele, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability for Neural TSP Solvers
von: Narad, Reuben, et al.
Veröffentlicht: (2025)
von: Narad, Reuben, et al.
Veröffentlicht: (2025)
Linear Causal Representation Learning by Topological Ordering, Pruning, and Disentanglement
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Reinforcement Learning Agents
von: Trim, Tristan, et al.
Veröffentlicht: (2024)
von: Trim, Tristan, et al.
Veröffentlicht: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
von: Palumbo, Nils, et al.
Veröffentlicht: (2024)
von: Palumbo, Nils, et al.
Veröffentlicht: (2024)
Score-based Causal Representation Learning: Linear and General Transformations
von: Varıcı, Burak, et al.
Veröffentlicht: (2024)
von: Varıcı, Burak, et al.
Veröffentlicht: (2024)
Causal Discovery of Linear Non-Gaussian Causal Models with Unobserved Confounding
von: Schkoda, Daniela, et al.
Veröffentlicht: (2024)
von: Schkoda, Daniela, et al.
Veröffentlicht: (2024)
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
von: Gonzalez, ML Nissen, et al.
Veröffentlicht: (2026)
von: Gonzalez, ML Nissen, et al.
Veröffentlicht: (2026)
InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
von: Gupta, Rohan, et al.
Veröffentlicht: (2024)
von: Gupta, Rohan, et al.
Veröffentlicht: (2024)
Linear Attention is Enough in Spatial-Temporal Forecasting
von: Ning, Xinyu
Veröffentlicht: (2024)
von: Ning, Xinyu
Veröffentlicht: (2024)
Mechanistic Interpretability for Transformer-based Time Series Classification
von: Kalnāre, Matīss, et al.
Veröffentlicht: (2025)
von: Kalnāre, Matīss, et al.
Veröffentlicht: (2025)
Causal Abstraction Learning based on the Semantic Embedding Principle
von: D'Acunto, Gabriele, et al.
Veröffentlicht: (2025)
von: D'Acunto, Gabriele, et al.
Veröffentlicht: (2025)
Linear Causal Representation Learning from Unknown Multi-node Interventions
von: Varıcı, Burak, et al.
Veröffentlicht: (2024)
von: Varıcı, Burak, et al.
Veröffentlicht: (2024)
Identifying General Mechanism Shifts in Linear Causal Representations
von: Chen, Tianyu, et al.
Veröffentlicht: (2024)
von: Chen, Tianyu, et al.
Veröffentlicht: (2024)
Finding Rule-Interpretable Non-Negative Data Representation
von: Mihelčić, Matej, et al.
Veröffentlicht: (2022)
von: Mihelčić, Matej, et al.
Veröffentlicht: (2022)
Multi-Domain Empirical Bayes for Linearly-Mixed Causal Representations
von: Wu, Bohan, et al.
Veröffentlicht: (2026)
von: Wu, Bohan, et al.
Veröffentlicht: (2026)
Geospatial Mechanistic Interpretability of Large Language Models
von: De Sabbata, Stef, et al.
Veröffentlicht: (2025)
von: De Sabbata, Stef, et al.
Veröffentlicht: (2025)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Interpretable Causal Representation Learning for Biological Data in the Pathway Space
von: de la Fuente, Jesus, et al.
Veröffentlicht: (2025)
von: de la Fuente, Jesus, et al.
Veröffentlicht: (2025)
Inferential Mechanics Part 1: Causal Mechanistic Theories of Machine Learning in Chemical Biology with Implications
von: Balabin, Ilya, et al.
Veröffentlicht: (2026)
von: Balabin, Ilya, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Causal Estimation of Memorisation Profiles
von: Lesci, Pietro, et al.
Veröffentlicht: (2024) -
Causal Abstraction Inference under Lossy Representations
von: Xia, Kevin, et al.
Veröffentlicht: (2025) -
Learning Causal Abstractions of Linear Structural Causal Models
von: Massidda, Riccardo, et al.
Veröffentlicht: (2024) -
Causal Estimation of Tokenisation Bias
von: Lesci, Pietro, et al.
Veröffentlicht: (2025) -
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
von: Minder, Julian, et al.
Veröffentlicht: (2025)