Causality is Key for Interpretability Claims to Generalise
Fuente:
arXiv
Salvato in:
| Autori principali: | Joshi, Shruti, Mueller, Aaron, Klindt, David, Brendel, Wieland, Reizinger, Patrik, Sridhar, Dhanya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations
di: Joshi, Shruti, et al.
Pubblicazione: (2026)
di: Joshi, Shruti, et al.
Pubblicazione: (2026)
Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
Estimating Treatment Effects with Independent Component Analysis
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
Cross-Entropy Is All You Need To Invert the Data Generating Process
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
An Interventional Perspective on Identifiability in Gaussian LTI Systems with Independent Component Analysis
di: Rajendran, Goutham, et al.
Pubblicazione: (2023)
di: Rajendran, Goutham, et al.
Pubblicazione: (2023)
Rule Extrapolation in Language Models: A Study of Compositional Generalization on OOD Prompts
di: Mészáros, Anna, et al.
Pubblicazione: (2024)
di: Mészáros, Anna, et al.
Pubblicazione: (2024)
Position: Understanding LLMs Requires More Than Statistical Generalization
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
di: Reizinger, Patrik, et al.
Pubblicazione: (2024)
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)
InfoNCE: Identifying the Gap Between Theory and Practice
di: Rusak, Evgenia, et al.
Pubblicazione: (2024)
di: Rusak, Evgenia, et al.
Pubblicazione: (2024)
From superposition to sparse codes: interpretable representations in neural networks
di: Klindt, David, et al.
Pubblicazione: (2025)
di: Klindt, David, et al.
Pubblicazione: (2025)
Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalisation
di: Pacela, Vitória Barin, et al.
Pubblicazione: (2026)
di: Pacela, Vitória Barin, et al.
Pubblicazione: (2026)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
di: Joshi, Shruti, et al.
Pubblicazione: (2025)
di: Joshi, Shruti, et al.
Pubblicazione: (2025)
Out-of-distribution Tests Reveal Compositionality in Chess Transformers
di: Mészáros, Anna, et al.
Pubblicazione: (2025)
di: Mészáros, Anna, et al.
Pubblicazione: (2025)
General Causal Imputation via Synthetic Interventions
di: Jiralerspong, Marco, et al.
Pubblicazione: (2024)
di: Jiralerspong, Marco, et al.
Pubblicazione: (2024)
The Landscape of Causal Discovery Data: Grounding Causal Discovery in Real-World Applications
di: Brouillard, Philippe, et al.
Pubblicazione: (2024)
di: Brouillard, Philippe, et al.
Pubblicazione: (2024)
Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks
di: Mueller, Aaron
Pubblicazione: (2024)
di: Mueller, Aaron
Pubblicazione: (2024)
Data Whitening Improves Sparse Autoencoder Learning
di: Saraswatula, Ashwin, et al.
Pubblicazione: (2025)
di: Saraswatula, Ashwin, et al.
Pubblicazione: (2025)
Causal Representation Learning in Temporal Data via Single-Parent Decoding
di: Brouillard, Philippe, et al.
Pubblicazione: (2024)
di: Brouillard, Philippe, et al.
Pubblicazione: (2024)
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
Identifiable Deep Generative Models via Sparse Decoding
di: Moran, Gemma E., et al.
Pubblicazione: (2021)
di: Moran, Gemma E., et al.
Pubblicazione: (2021)
Demystifying amortized causal discovery with transformers
di: Montagna, Francesco, et al.
Pubblicazione: (2024)
di: Montagna, Francesco, et al.
Pubblicazione: (2024)
Superposition disentanglement of neural representations reveals hidden alignment
di: Longon, André, et al.
Pubblicazione: (2025)
di: Longon, André, et al.
Pubblicazione: (2025)
The Role of Causal Features in Strategic Classification for Robustness and Alignment
di: Gois, Antonio, et al.
Pubblicazione: (2026)
di: Gois, Antonio, et al.
Pubblicazione: (2026)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
di: Marks, Samuel, et al.
Pubblicazione: (2024)
di: Marks, Samuel, et al.
Pubblicazione: (2024)
In-Context Learning Can Re-learn Forbidden Tasks
di: Xhonneux, Sophie, et al.
Pubblicazione: (2024)
di: Xhonneux, Sophie, et al.
Pubblicazione: (2024)
Provable Compositional Generalization for Object-Centric Learning
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2023)
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2023)
Sparsity regularization via tree-structured environments for disentangled representations
di: Layne, Elliot, et al.
Pubblicazione: (2024)
di: Layne, Elliot, et al.
Pubblicazione: (2024)
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
di: O'Neill, Charles, et al.
Pubblicazione: (2024)
di: O'Neill, Charles, et al.
Pubblicazione: (2024)
Occam's Razor for Self Supervised Learning: What is Sufficient to Learn Good Representations?
di: Ibrahim, Mark, et al.
Pubblicazione: (2024)
di: Ibrahim, Mark, et al.
Pubblicazione: (2024)
LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models
di: Li, Fanfei, et al.
Pubblicazione: (2025)
di: Li, Fanfei, et al.
Pubblicazione: (2025)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
Generation is Required for Data-Efficient Perception
di: Brady, Jack, et al.
Pubblicazione: (2025)
di: Brady, Jack, et al.
Pubblicazione: (2025)
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2023)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2023)
CSQL: Mapping Documents into Causal Databases
di: Mahadevan, Sridhar
Pubblicazione: (2026)
di: Mahadevan, Sridhar
Pubblicazione: (2026)
The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
di: Mueller, Aaron, et al.
Pubblicazione: (2024)
di: Mueller, Aaron, et al.
Pubblicazione: (2024)
Stress-Testing Causal Claims via Cardinality Repairs
di: Gabbay, Yarden, et al.
Pubblicazione: (2025)
di: Gabbay, Yarden, et al.
Pubblicazione: (2025)
Interaction Asymmetry: A General Principle for Learning Composable Abstractions
di: Brady, Jack, et al.
Pubblicazione: (2024)
di: Brady, Jack, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations
di: Joshi, Shruti, et al.
Pubblicazione: (2026) -
Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research
di: Reizinger, Patrik, et al.
Pubblicazione: (2025) -
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
di: Mueller, Aaron, et al.
Pubblicazione: (2025) -
Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning
di: Reizinger, Patrik, et al.
Pubblicazione: (2024) -
Estimating Treatment Effects with Independent Component Analysis
di: Reizinger, Patrik, et al.
Pubblicazione: (2025)