When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Gonzalez, ML Nissen, Albuquerque, Melwina, Wroe, Laurence, Cohen, Jacob Meyer, Smith, Logan Riggs, Dooms, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Mechanistic to Compositional Interpretability
by: Gauderis, Ward, et al.
Published: (2026)
by: Gauderis, Ward, et al.
Published: (2026)
Compositionality Unlocks Deep Interpretable Models
by: Dooms, Thomas, et al.
Published: (2025)
by: Dooms, Thomas, et al.
Published: (2025)
Finding Manifolds With Bilinear Autoencoders
by: Dooms, Thomas, et al.
Published: (2025)
by: Dooms, Thomas, et al.
Published: (2025)
Tokenized SAEs: Disentangling SAE Reconstructions
by: Dooms, Thomas, et al.
Published: (2025)
by: Dooms, Thomas, et al.
Published: (2025)
Decomposition of Small Transformer Models
by: Christensen, Casper L., et al.
Published: (2025)
by: Christensen, Casper L., et al.
Published: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
by: Engels, Joshua, et al.
Published: (2024)
by: Engels, Joshua, et al.
Published: (2024)
Fractals made Practical: Denoising Diffusion as Partitioned Iterated Function Systems
by: Dooms, Ann
Published: (2026)
by: Dooms, Ann
Published: (2026)
Weight-based Decomposition: A Case for Bilinear MLPs
by: Pearce, Michael T., et al.
Published: (2024)
by: Pearce, Michael T., et al.
Published: (2024)
Finite Basis Physics-Informed Neural Networks (FBPINNs): a scalable domain decomposition approach for solving differential equations
by: Moseley, Ben, et al.
Published: (2021)
by: Moseley, Ben, et al.
Published: (2021)
Bilinear autoencoders find interpretable manifolds
by: Dooms, Thomas, et al.
Published: (2026)
by: Dooms, Thomas, et al.
Published: (2026)
When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability
by: Ronge, Raphael, et al.
Published: (2026)
by: Ronge, Raphael, et al.
Published: (2026)
Exemplar Partitioning for Mechanistic Interpretability
by: Rumbelow, Jessica
Published: (2026)
by: Rumbelow, Jessica
Published: (2026)
Open Problems in Mechanistic Interpretability
by: Sharkey, Lee, et al.
Published: (2025)
by: Sharkey, Lee, et al.
Published: (2025)
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
by: Sun, Alan, et al.
Published: (2026)
by: Sun, Alan, et al.
Published: (2026)
Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
by: Kowalska, Bianka, et al.
Published: (2025)
by: Kowalska, Bianka, et al.
Published: (2025)
Bilinear MLPs enable weight-based mechanistic interpretability
by: Pearce, Michael T., et al.
Published: (2024)
by: Pearce, Michael T., et al.
Published: (2024)
Interpretable Tensor Fusion
by: Varshneya, Saurabh, et al.
Published: (2024)
by: Varshneya, Saurabh, et al.
Published: (2024)
Mechanistic Interpretability for Neural TSP Solvers
by: Narad, Reuben, et al.
Published: (2025)
by: Narad, Reuben, et al.
Published: (2025)
Mechanistic Interpretability of Reinforcement Learning Agents
by: Trim, Tristan, et al.
Published: (2024)
by: Trim, Tristan, et al.
Published: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
by: Palumbo, Nils, et al.
Published: (2024)
by: Palumbo, Nils, et al.
Published: (2024)
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
by: Sutter, Denis, et al.
Published: (2025)
by: Sutter, Denis, et al.
Published: (2025)
InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
by: Gupta, Rohan, et al.
Published: (2024)
by: Gupta, Rohan, et al.
Published: (2024)
Mechanistic Interpretability for Transformer-based Time Series Classification
by: Kalnāre, Matīss, et al.
Published: (2025)
by: Kalnāre, Matīss, et al.
Published: (2025)
Geospatial Mechanistic Interpretability of Large Language Models
by: De Sabbata, Stef, et al.
Published: (2025)
by: De Sabbata, Stef, et al.
Published: (2025)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
by: Bushnaq, Lucius, et al.
Published: (2024)
by: Bushnaq, Lucius, et al.
Published: (2024)
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
When Fusion Helps and When It Breaks: View-Aligned Robustness in Same-Source Financial Imaging
by: Ma, Rui
Published: (2026)
by: Ma, Rui
Published: (2026)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
by: Saini, Harshvardhan, et al.
Published: (2026)
by: Saini, Harshvardhan, et al.
Published: (2026)
Interpretable Bayesian Tensor Network Kernel Machines with Automatic Rank and Feature Selection
by: Kilic, Afra, et al.
Published: (2025)
by: Kilic, Afra, et al.
Published: (2025)
Challenges in Mechanistically Interpreting Model Representations
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
Mechanistic Interpretability of Binary and Ternary Transformers
by: Li, Jason
Published: (2024)
by: Li, Jason
Published: (2024)
Mechanistic Interpretability of RNNs emulating Hidden Markov Models
by: Torre, Elia, et al.
Published: (2025)
by: Torre, Elia, et al.
Published: (2025)
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
by: Miller, Ryan J., et al.
Published: (2025)
by: Miller, Ryan J., et al.
Published: (2025)
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
by: Winninger, Thomas, et al.
Published: (2025)
by: Winninger, Thomas, et al.
Published: (2025)
Compact Proofs of Model Performance via Mechanistic Interpretability
by: Gross, Jason, et al.
Published: (2024)
by: Gross, Jason, et al.
Published: (2024)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
by: Erdogan, Ege, et al.
Published: (2025)
by: Erdogan, Ege, et al.
Published: (2025)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
by: He, Jesse, et al.
Published: (2026)
by: He, Jesse, et al.
Published: (2026)
Triangulation as an Acceptance Rule for Multilingual Mechanistic Interpretability
by: Long, Yanan
Published: (2025)
by: Long, Yanan
Published: (2025)
SIC: Similarity-Based Interpretable Image Classification with Neural Networks
by: Wolf, Tom Nuno, et al.
Published: (2025)
by: Wolf, Tom Nuno, et al.
Published: (2025)
Similar Items
-
From Mechanistic to Compositional Interpretability
by: Gauderis, Ward, et al.
Published: (2026) -
Compositionality Unlocks Deep Interpretable Models
by: Dooms, Thomas, et al.
Published: (2025) -
Finding Manifolds With Bilinear Autoencoders
by: Dooms, Thomas, et al.
Published: (2025) -
Tokenized SAEs: Disentangling SAE Reconstructions
by: Dooms, Thomas, et al.
Published: (2025) -
Decomposition of Small Transformer Models
by: Christensen, Casper L., et al.
Published: (2025)