When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
Fuente:
arXiv
Guardado en:
| Autores principales: | Gonzalez, ML Nissen, Albuquerque, Melwina, Wroe, Laurence, Cohen, Jacob Meyer, Smith, Logan Riggs, Dooms, Thomas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Mechanistic to Compositional Interpretability
por: Gauderis, Ward, et al.
Publicado: (2026)
por: Gauderis, Ward, et al.
Publicado: (2026)
Compositionality Unlocks Deep Interpretable Models
por: Dooms, Thomas, et al.
Publicado: (2025)
por: Dooms, Thomas, et al.
Publicado: (2025)
Finding Manifolds With Bilinear Autoencoders
por: Dooms, Thomas, et al.
Publicado: (2025)
por: Dooms, Thomas, et al.
Publicado: (2025)
Tokenized SAEs: Disentangling SAE Reconstructions
por: Dooms, Thomas, et al.
Publicado: (2025)
por: Dooms, Thomas, et al.
Publicado: (2025)
Decomposition of Small Transformer Models
por: Christensen, Casper L., et al.
Publicado: (2025)
por: Christensen, Casper L., et al.
Publicado: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
por: Engels, Joshua, et al.
Publicado: (2024)
por: Engels, Joshua, et al.
Publicado: (2024)
Fractals made Practical: Denoising Diffusion as Partitioned Iterated Function Systems
por: Dooms, Ann
Publicado: (2026)
por: Dooms, Ann
Publicado: (2026)
Weight-based Decomposition: A Case for Bilinear MLPs
por: Pearce, Michael T., et al.
Publicado: (2024)
por: Pearce, Michael T., et al.
Publicado: (2024)
Finite Basis Physics-Informed Neural Networks (FBPINNs): a scalable domain decomposition approach for solving differential equations
por: Moseley, Ben, et al.
Publicado: (2021)
por: Moseley, Ben, et al.
Publicado: (2021)
Bilinear autoencoders find interpretable manifolds
por: Dooms, Thomas, et al.
Publicado: (2026)
por: Dooms, Thomas, et al.
Publicado: (2026)
When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability
por: Ronge, Raphael, et al.
Publicado: (2026)
por: Ronge, Raphael, et al.
Publicado: (2026)
Exemplar Partitioning for Mechanistic Interpretability
por: Rumbelow, Jessica
Publicado: (2026)
por: Rumbelow, Jessica
Publicado: (2026)
Open Problems in Mechanistic Interpretability
por: Sharkey, Lee, et al.
Publicado: (2025)
por: Sharkey, Lee, et al.
Publicado: (2025)
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
por: Sun, Alan, et al.
Publicado: (2026)
por: Sun, Alan, et al.
Publicado: (2026)
Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
por: Kowalska, Bianka, et al.
Publicado: (2025)
por: Kowalska, Bianka, et al.
Publicado: (2025)
Bilinear MLPs enable weight-based mechanistic interpretability
por: Pearce, Michael T., et al.
Publicado: (2024)
por: Pearce, Michael T., et al.
Publicado: (2024)
Interpretable Tensor Fusion
por: Varshneya, Saurabh, et al.
Publicado: (2024)
por: Varshneya, Saurabh, et al.
Publicado: (2024)
Mechanistic Interpretability for Neural TSP Solvers
por: Narad, Reuben, et al.
Publicado: (2025)
por: Narad, Reuben, et al.
Publicado: (2025)
Mechanistic Interpretability of Reinforcement Learning Agents
por: Trim, Tristan, et al.
Publicado: (2024)
por: Trim, Tristan, et al.
Publicado: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
por: Palumbo, Nils, et al.
Publicado: (2024)
por: Palumbo, Nils, et al.
Publicado: (2024)
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
por: Sutter, Denis, et al.
Publicado: (2025)
por: Sutter, Denis, et al.
Publicado: (2025)
InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
por: Gupta, Rohan, et al.
Publicado: (2024)
por: Gupta, Rohan, et al.
Publicado: (2024)
Mechanistic Interpretability for Transformer-based Time Series Classification
por: Kalnāre, Matīss, et al.
Publicado: (2025)
por: Kalnāre, Matīss, et al.
Publicado: (2025)
Geospatial Mechanistic Interpretability of Large Language Models
por: De Sabbata, Stef, et al.
Publicado: (2025)
por: De Sabbata, Stef, et al.
Publicado: (2025)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
por: Bushnaq, Lucius, et al.
Publicado: (2024)
por: Bushnaq, Lucius, et al.
Publicado: (2024)
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
por: Heap, Thomas, et al.
Publicado: (2025)
por: Heap, Thomas, et al.
Publicado: (2025)
When Fusion Helps and When It Breaks: View-Aligned Robustness in Same-Source Financial Imaging
por: Ma, Rui
Publicado: (2026)
por: Ma, Rui
Publicado: (2026)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
por: Saini, Harshvardhan, et al.
Publicado: (2026)
por: Saini, Harshvardhan, et al.
Publicado: (2026)
Interpretable Bayesian Tensor Network Kernel Machines with Automatic Rank and Feature Selection
por: Kilic, Afra, et al.
Publicado: (2025)
por: Kilic, Afra, et al.
Publicado: (2025)
Challenges in Mechanistically Interpreting Model Representations
por: Golechha, Satvik, et al.
Publicado: (2024)
por: Golechha, Satvik, et al.
Publicado: (2024)
Mechanistic Interpretability of Binary and Ternary Transformers
por: Li, Jason
Publicado: (2024)
por: Li, Jason
Publicado: (2024)
Mechanistic Interpretability of RNNs emulating Hidden Markov Models
por: Torre, Elia, et al.
Publicado: (2025)
por: Torre, Elia, et al.
Publicado: (2025)
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
por: Miller, Ryan J., et al.
Publicado: (2025)
por: Miller, Ryan J., et al.
Publicado: (2025)
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
por: Winninger, Thomas, et al.
Publicado: (2025)
por: Winninger, Thomas, et al.
Publicado: (2025)
Compact Proofs of Model Performance via Mechanistic Interpretability
por: Gross, Jason, et al.
Publicado: (2024)
por: Gross, Jason, et al.
Publicado: (2024)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
por: Song, Xiangchen, et al.
Publicado: (2025)
por: Song, Xiangchen, et al.
Publicado: (2025)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
por: Erdogan, Ege, et al.
Publicado: (2025)
por: Erdogan, Ege, et al.
Publicado: (2025)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
por: He, Jesse, et al.
Publicado: (2026)
por: He, Jesse, et al.
Publicado: (2026)
Triangulation as an Acceptance Rule for Multilingual Mechanistic Interpretability
por: Long, Yanan
Publicado: (2025)
por: Long, Yanan
Publicado: (2025)
SIC: Similarity-Based Interpretable Image Classification with Neural Networks
por: Wolf, Tom Nuno, et al.
Publicado: (2025)
por: Wolf, Tom Nuno, et al.
Publicado: (2025)
Ejemplares similares
-
From Mechanistic to Compositional Interpretability
por: Gauderis, Ward, et al.
Publicado: (2026) -
Compositionality Unlocks Deep Interpretable Models
por: Dooms, Thomas, et al.
Publicado: (2025) -
Finding Manifolds With Bilinear Autoencoders
por: Dooms, Thomas, et al.
Publicado: (2025) -
Tokenized SAEs: Disentangling SAE Reconstructions
por: Dooms, Thomas, et al.
Publicado: (2025) -
Decomposition of Small Transformer Models
por: Christensen, Casper L., et al.
Publicado: (2025)