Exemplar Partitioning for Mechanistic Interpretability
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Rumbelow, Jessica |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Open Problems in Mechanistic Interpretability
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability for Neural TSP Solvers
von: Narad, Reuben, et al.
Veröffentlicht: (2025)
von: Narad, Reuben, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Reinforcement Learning Agents
von: Trim, Tristan, et al.
Veröffentlicht: (2024)
von: Trim, Tristan, et al.
Veröffentlicht: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
von: Palumbo, Nils, et al.
Veröffentlicht: (2024)
von: Palumbo, Nils, et al.
Veröffentlicht: (2024)
GnnXemplar: Exemplars to Explanations -- Natural Language Rules for Global GNN Interpretability
von: Armgaan, Burouj, et al.
Veröffentlicht: (2025)
von: Armgaan, Burouj, et al.
Veröffentlicht: (2025)
Geospatial Mechanistic Interpretability of Large Language Models
von: De Sabbata, Stef, et al.
Veröffentlicht: (2025)
von: De Sabbata, Stef, et al.
Veröffentlicht: (2025)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
The SaTML '24 CNN Interpretability Competition: New Innovations for Concept-Level Interpretability
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
Challenges in Mechanistically Interpreting Model Representations
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of Binary and Ternary Transformers
von: Li, Jason
Veröffentlicht: (2024)
von: Li, Jason
Veröffentlicht: (2024)
Mechanistic Interpretability of RNNs emulating Hidden Markov Models
von: Torre, Elia, et al.
Veröffentlicht: (2025)
von: Torre, Elia, et al.
Veröffentlicht: (2025)
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
von: Miller, Ryan J., et al.
Veröffentlicht: (2025)
von: Miller, Ryan J., et al.
Veröffentlicht: (2025)
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
von: Gonzalez, ML Nissen, et al.
Veröffentlicht: (2026)
von: Gonzalez, ML Nissen, et al.
Veröffentlicht: (2026)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
von: Erdogan, Ege, et al.
Veröffentlicht: (2025)
von: Erdogan, Ege, et al.
Veröffentlicht: (2025)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
von: He, Jesse, et al.
Veröffentlicht: (2026)
von: He, Jesse, et al.
Veröffentlicht: (2026)
Triangulation as an Acceptance Rule for Multilingual Mechanistic Interpretability
von: Long, Yanan
Veröffentlicht: (2025)
von: Long, Yanan
Veröffentlicht: (2025)
Exemplar-Free Continual Learning for State Space Models
von: Lee, Isaac Ning, et al.
Veröffentlicht: (2025)
von: Lee, Isaac Ning, et al.
Veröffentlicht: (2025)
EXPLORA: Efficient Exemplar Subset Selection for Complex Reasoning
von: Purohit, Kiran, et al.
Veröffentlicht: (2024)
von: Purohit, Kiran, et al.
Veröffentlicht: (2024)
MIB: A Mechanistic Interpretability Benchmark
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
Putting a Face to Forgetting: Continual Learning meets Mechanistic Interpretability
von: Masip, Sergi, et al.
Veröffentlicht: (2026)
von: Masip, Sergi, et al.
Veröffentlicht: (2026)
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
von: Sutter, Denis, et al.
Veröffentlicht: (2025)
von: Sutter, Denis, et al.
Veröffentlicht: (2025)
InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
von: Gupta, Rohan, et al.
Veröffentlicht: (2024)
von: Gupta, Rohan, et al.
Veröffentlicht: (2024)
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
von: Kowalska, Bianka, et al.
Veröffentlicht: (2025)
von: Kowalska, Bianka, et al.
Veröffentlicht: (2025)
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
von: Sun, Alan, et al.
Veröffentlicht: (2026)
von: Sun, Alan, et al.
Veröffentlicht: (2026)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
von: Dumas, Clément
Veröffentlicht: (2025)
von: Dumas, Clément
Veröffentlicht: (2025)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
von: Yan, Ge, et al.
Veröffentlicht: (2025)
von: Yan, Ge, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability for Transformer-based Time Series Classification
von: Kalnāre, Matīss, et al.
Veröffentlicht: (2025)
von: Kalnāre, Matīss, et al.
Veröffentlicht: (2025)
OceanCBM: A Concept Bottleneck Model for Mechanistic Interpretability in Ocean Forecasting
von: Suri, Sanah, et al.
Veröffentlicht: (2026)
von: Suri, Sanah, et al.
Veröffentlicht: (2026)
MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
von: Khadka, Barsat
Veröffentlicht: (2026)
von: Khadka, Barsat
Veröffentlicht: (2026)
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
von: Roy, Dip, et al.
Veröffentlicht: (2025)
von: Roy, Dip, et al.
Veröffentlicht: (2025)
Exemplar-condensed Federated Class-incremental Learning
von: Sun, Rui, et al.
Veröffentlicht: (2024)
von: Sun, Rui, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
von: Tolooshams, Bahareh, et al.
Veröffentlicht: (2025)
von: Tolooshams, Bahareh, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
reward-lens: A Mechanistic Interpretability Library for Reward Models
von: Nadaf, Mohammed Suhail B
Veröffentlicht: (2026)
von: Nadaf, Mohammed Suhail B
Veröffentlicht: (2026)
Compact Proofs of Model Performance via Mechanistic Interpretability
von: Gross, Jason, et al.
Veröffentlicht: (2024)
von: Gross, Jason, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
von: El, Batu, et al.
Veröffentlicht: (2025)
von: El, Batu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Open Problems in Mechanistic Interpretability
von: Sharkey, Lee, et al.
Veröffentlicht: (2025) -
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026) -
Mechanistic Interpretability for Neural TSP Solvers
von: Narad, Reuben, et al.
Veröffentlicht: (2025) -
Mechanistic Interpretability of Reinforcement Learning Agents
von: Trim, Tristan, et al.
Veröffentlicht: (2024) -
Validating Mechanistic Interpretations: An Axiomatic Approach
von: Palumbo, Nils, et al.
Veröffentlicht: (2024)