From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
Fuente:
arXiv
Saved in:
| Main Authors: | Mueller, Aaron, Lee, Andrew, Joshi, Shruti, Lubana, Ekdeep Singh, Sridhar, Dhanya, Reizinger, Patrik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causality is Key for Interpretability Claims to Generalise
by: Joshi, Shruti, et al.
Published: (2026)
by: Joshi, Shruti, et al.
Published: (2026)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
by: Joshi, Shruti, et al.
Published: (2025)
by: Joshi, Shruti, et al.
Published: (2025)
Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations
by: Joshi, Shruti, et al.
Published: (2026)
by: Joshi, Shruti, et al.
Published: (2026)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
by: Jaipersaud, Brandon, et al.
Published: (2025)
by: Jaipersaud, Brandon, et al.
Published: (2025)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
by: Zur, Amir, et al.
Published: (2025)
by: Zur, Amir, et al.
Published: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
by: Pres, Itamar, et al.
Published: (2024)
by: Pres, Itamar, et al.
Published: (2024)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
ICLR: In-Context Learning of Representations
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
In-Context Learning Dynamics with Random Binary Sequences
by: Bigelow, Eric J., et al.
Published: (2023)
by: Bigelow, Eric J., et al.
Published: (2023)
Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks
by: Mueller, Aaron
Published: (2024)
by: Mueller, Aaron
Published: (2024)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
by: Hindupur, Sai Sumedh R., et al.
Published: (2025)
by: Hindupur, Sai Sumedh R., et al.
Published: (2025)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
by: Gopalani, Pulkit, et al.
Published: (2024)
by: Gopalani, Pulkit, et al.
Published: (2024)
Rule Extrapolation in Language Models: A Study of Compositional Generalization on OOD Prompts
by: Mészáros, Anna, et al.
Published: (2024)
by: Mészáros, Anna, et al.
Published: (2024)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
by: Bigelow, Eric, et al.
Published: (2025)
by: Bigelow, Eric, et al.
Published: (2025)
Emergence of Hierarchical Emotion Organization in Large Language Models
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
by: Jain, Samyak, et al.
Published: (2024)
by: Jain, Samyak, et al.
Published: (2024)
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
by: Ramesh, Rahul, et al.
Published: (2023)
by: Ramesh, Rahul, et al.
Published: (2023)
The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
by: Sarfati, Raphaël, et al.
Published: (2026)
by: Sarfati, Raphaël, et al.
Published: (2026)
Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
by: Feucht, Sheridan, et al.
Published: (2026)
by: Feucht, Sheridan, et al.
Published: (2026)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
by: Bigelow, Eric, et al.
Published: (2026)
by: Bigelow, Eric, et al.
Published: (2026)
Swing-by Dynamics in Concept Learning and Compositional Generalization
by: Yang, Yongyi, et al.
Published: (2024)
by: Yang, Yongyi, et al.
Published: (2024)
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025)
by: Costa, Valérie, et al.
Published: (2025)
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025)
by: Costa, Valérie, et al.
Published: (2025)
An Interventional Perspective on Identifiability in Gaussian LTI Systems with Independent Component Analysis
by: Rajendran, Goutham, et al.
Published: (2023)
by: Rajendran, Goutham, et al.
Published: (2023)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
by: Bohacek, Matyas, et al.
Published: (2025)
by: Bohacek, Matyas, et al.
Published: (2025)
Analyzing (In)Abilities of SAEs via Formal Languages
by: Menon, Abhinav, et al.
Published: (2024)
by: Menon, Abhinav, et al.
Published: (2024)
Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research
by: Reizinger, Patrik, et al.
Published: (2025)
by: Reizinger, Patrik, et al.
Published: (2025)
Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
by: Brinkmann, Jannik, et al.
Published: (2025)
by: Brinkmann, Jannik, et al.
Published: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
by: Bhalla, Usha, et al.
Published: (2026)
by: Bhalla, Usha, et al.
Published: (2026)
When Interpretability Becomes a Liability: Adversarial Attacks on CBM Concept Layers
by: Sridhar, Aditya
Published: (2026)
by: Sridhar, Aditya
Published: (2026)
Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts
by: Peng, Kenny, et al.
Published: (2025)
by: Peng, Kenny, et al.
Published: (2025)
Concept-Based Interpretability for Toxicity Detection
by: Garg, Samarth, et al.
Published: (2025)
by: Garg, Samarth, et al.
Published: (2025)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
by: Huang, Jing, et al.
Published: (2024)
by: Huang, Jing, et al.
Published: (2024)
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
by: Lubana, Ekdeep Singh, et al.
Published: (2024)
by: Lubana, Ekdeep Singh, et al.
Published: (2024)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
by: Okawa, Maya, et al.
Published: (2023)
by: Okawa, Maya, et al.
Published: (2023)
Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning
by: Reizinger, Patrik, et al.
Published: (2024)
by: Reizinger, Patrik, et al.
Published: (2024)
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
by: Fel, Thomas, et al.
Published: (2025)
by: Fel, Thomas, et al.
Published: (2025)
Sparse Graph Representations for Procedural Instructional Documents
by: Singh, Shruti, et al.
Published: (2024)
by: Singh, Shruti, et al.
Published: (2024)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
Similar Items
-
Causality is Key for Interpretability Claims to Generalise
by: Joshi, Shruti, et al.
Published: (2026) -
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
by: Joshi, Shruti, et al.
Published: (2025) -
Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations
by: Joshi, Shruti, et al.
Published: (2026) -
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
by: Jaipersaud, Brandon, et al.
Published: (2025) -
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
by: Zur, Amir, et al.
Published: (2025)