From superposition to sparse codes: interpretable representations in neural networks
Fuente:
arXiv
Saved in:
| Main Authors: | Klindt, David, O'Neill, Charles, Reizinger, Patrik, Maurer, Harald, Miolane, Nina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
by: O'Neill, Charles, et al.
Published: (2024)
by: O'Neill, Charles, et al.
Published: (2024)
Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research
by: Reizinger, Patrik, et al.
Published: (2025)
by: Reizinger, Patrik, et al.
Published: (2025)
Causality is Key for Interpretability Claims to Generalise
by: Joshi, Shruti, et al.
Published: (2026)
by: Joshi, Shruti, et al.
Published: (2026)
Superposition disentanglement of neural representations reveals hidden alignment
by: Longon, André, et al.
Published: (2025)
by: Longon, André, et al.
Published: (2025)
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
by: O'Neill, Charles
Published: (2025)
by: O'Neill, Charles
Published: (2025)
Cross-Entropy Is All You Need To Invert the Data Generating Process
by: Reizinger, Patrik, et al.
Published: (2024)
by: Reizinger, Patrik, et al.
Published: (2024)
Out-of-distribution Tests Reveal Compositionality in Chess Transformers
by: Mészáros, Anna, et al.
Published: (2025)
by: Mészáros, Anna, et al.
Published: (2025)
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
by: O'Neill, Charles, et al.
Published: (2024)
by: O'Neill, Charles, et al.
Published: (2024)
Type 2 Tobit Sample Selection Models with Bayesian Additive Regression Trees
by: O'Neill, Eoghan
Published: (2025)
by: O'Neill, Eoghan
Published: (2025)
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
by: O'Neill, Charles, et al.
Published: (2025)
by: O'Neill, Charles, et al.
Published: (2025)
Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity
by: Miller, Jack, et al.
Published: (2023)
by: Miller, Jack, et al.
Published: (2023)
Estimating Treatment Effects with Independent Component Analysis
by: Reizinger, Patrik, et al.
Published: (2025)
by: Reizinger, Patrik, et al.
Published: (2025)
A General Framework for Robust G-Invariance in G-Equivariant Networks
by: Sanborn, Sophia, et al.
Published: (2023)
by: Sanborn, Sophia, et al.
Published: (2023)
Data Whitening Improves Sparse Autoencoder Learning
by: Saraswatula, Ashwin, et al.
Published: (2025)
by: Saraswatula, Ashwin, et al.
Published: (2025)
Motif distribution and function of sparse deep neural networks
by: Zahn, Olivia T., et al.
Published: (2024)
by: Zahn, Olivia T., et al.
Published: (2024)
Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning
by: Reizinger, Patrik, et al.
Published: (2024)
by: Reizinger, Patrik, et al.
Published: (2024)
Rule Extrapolation in Language Models: A Study of Compositional Generalization on OOD Prompts
by: Mészáros, Anna, et al.
Published: (2024)
by: Mészáros, Anna, et al.
Published: (2024)
An Interventional Perspective on Identifiability in Gaussian LTI Systems with Independent Component Analysis
by: Rajendran, Goutham, et al.
Published: (2023)
by: Rajendran, Goutham, et al.
Published: (2023)
Disentangling Dense Embeddings with Sparse Autoencoders
by: O'Neill, Charles, et al.
Published: (2024)
by: O'Neill, Charles, et al.
Published: (2024)
Biology-informed neural networks learn nonlinear representations from omics data to improve genomic prediction and interpretability
by: Kontolati, Katiana, et al.
Published: (2025)
by: Kontolati, Katiana, et al.
Published: (2025)
Embedding interpretable $\ell_1$-regression into neural networks for uncovering temporal structure in cell imaging
by: Kabus, Fabian, et al.
Published: (2026)
by: Kabus, Fabian, et al.
Published: (2026)
Position: Understanding LLMs Requires More Than Statistical Generalization
by: Reizinger, Patrik, et al.
Published: (2024)
by: Reizinger, Patrik, et al.
Published: (2024)
Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations
by: Joshi, Shruti, et al.
Published: (2026)
by: Joshi, Shruti, et al.
Published: (2026)
Sketching the Heat Kernel: Using Gaussian Processes to Embed Data
by: Gilbert, Anna C., et al.
Published: (2024)
by: Gilbert, Anna C., et al.
Published: (2024)
When Machine Learning Gets Personal: Evaluating Prediction and Explanation
by: Cornelis, Louisa, et al.
Published: (2025)
by: Cornelis, Louisa, et al.
Published: (2025)
bispectrum: Selective $G$-Bispectra Made Practical
by: Mathe, Johan, et al.
Published: (2026)
by: Mathe, Johan, et al.
Published: (2026)
Architectures of Topological Deep Learning: A Survey of Message-Passing Topological Neural Networks
by: Papillon, Mathilde, et al.
Published: (2023)
by: Papillon, Mathilde, et al.
Published: (2023)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Measuring Sharpness in Grokking
by: Miller, Jack, et al.
Published: (2024)
by: Miller, Jack, et al.
Published: (2024)
Modelling the Doughnut of social and planetary boundaries with frugal machine learning
by: Vrizzi, Stefano, et al.
Published: (2025)
by: Vrizzi, Stefano, et al.
Published: (2025)
A mechanistically interpretable neural network for regulatory genomics
by: Tseng, Alex M., et al.
Published: (2024)
by: Tseng, Alex M., et al.
Published: (2024)
Tensorization of neural networks for improved privacy and interpretability
by: Monturiol, José Ramón Pareja, et al.
Published: (2025)
by: Monturiol, José Ramón Pareja, et al.
Published: (2025)
Bayesian neural networks with interpretable priors from Mercer kernels
by: Alberts, Alex, et al.
Published: (2025)
by: Alberts, Alex, et al.
Published: (2025)
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
by: Reizinger, Patrik, et al.
Published: (2025)
by: Reizinger, Patrik, et al.
Published: (2025)
Topological safeguard for evasion attack interpreting the neural networks' behavior
by: Echeberria-Barrio, Xabier, et al.
Published: (2024)
by: Echeberria-Barrio, Xabier, et al.
Published: (2024)
Robustness in sparse artificial neural networks trained with adaptive topology
by: Sulyok, Bendegúz, et al.
Published: (2026)
by: Sulyok, Bendegúz, et al.
Published: (2026)
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
by: O'Neill, Charles, et al.
Published: (2025)
by: O'Neill, Charles, et al.
Published: (2025)
PVNet: A LRCN Architecture for Spatio-Temporal Photovoltaic PowerForecasting from Numerical Weather Prediction
by: Mathe, Johan, et al.
Published: (2019)
by: Mathe, Johan, et al.
Published: (2019)
TopoTune : A Framework for Generalized Combinatorial Complex Neural Networks
by: Papillon, Mathilde, et al.
Published: (2024)
by: Papillon, Mathilde, et al.
Published: (2024)
Spike-and-slab shrinkage priors for structurally sparse Bayesian neural networks
by: Jantre, Sanket, et al.
Published: (2023)
by: Jantre, Sanket, et al.
Published: (2023)
Similar Items
-
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
by: O'Neill, Charles, et al.
Published: (2024) -
Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research
by: Reizinger, Patrik, et al.
Published: (2025) -
Causality is Key for Interpretability Claims to Generalise
by: Joshi, Shruti, et al.
Published: (2026) -
Superposition disentanglement of neural representations reveals hidden alignment
by: Longon, André, et al.
Published: (2025) -
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
by: O'Neill, Charles
Published: (2025)