Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
Fuente:
arXiv
Guardado en:
| Autores principales: | O'Neill, Charles, Jayasekara, Mudith, Kirkby, Max |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
por: O'Neill, Charles, et al.
Publicado: (2025)
por: O'Neill, Charles, et al.
Publicado: (2025)
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
por: O'Neill, Charles, et al.
Publicado: (2024)
por: O'Neill, Charles, et al.
Publicado: (2024)
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
por: O'Neill, Charles, et al.
Publicado: (2024)
por: O'Neill, Charles, et al.
Publicado: (2024)
Disentangling Dense Embeddings with Sparse Autoencoders
por: O'Neill, Charles, et al.
Publicado: (2024)
por: O'Neill, Charles, et al.
Publicado: (2024)
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
por: O'Neill, Charles
Publicado: (2025)
por: O'Neill, Charles
Publicado: (2025)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
por: Tolooshams, Bahareh, et al.
Publicado: (2025)
por: Tolooshams, Bahareh, et al.
Publicado: (2025)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
por: Erdogan, Ege, et al.
Publicado: (2025)
por: Erdogan, Ege, et al.
Publicado: (2025)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
por: Tahimic, Kriz, et al.
Publicado: (2025)
por: Tahimic, Kriz, et al.
Publicado: (2025)
Type 2 Tobit Sample Selection Models with Bayesian Additive Regression Trees
por: O'Neill, Eoghan
Publicado: (2025)
por: O'Neill, Eoghan
Publicado: (2025)
Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
por: Lehn-Schiøler, William, et al.
Publicado: (2026)
por: Lehn-Schiøler, William, et al.
Publicado: (2026)
Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity
por: Miller, Jack, et al.
Publicado: (2023)
por: Miller, Jack, et al.
Publicado: (2023)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
por: Wang, Xu, et al.
Publicado: (2026)
por: Wang, Xu, et al.
Publicado: (2026)
Transcoders Beat Sparse Autoencoders for Interpretability
por: Paulo, Gonçalo, et al.
Publicado: (2025)
por: Paulo, Gonçalo, et al.
Publicado: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
por: Engels, Joshua, et al.
Publicado: (2024)
por: Engels, Joshua, et al.
Publicado: (2024)
Interpretable Reward Model via Sparse Autoencoder
por: Zhang, Shuyi, et al.
Publicado: (2025)
por: Zhang, Shuyi, et al.
Publicado: (2025)
Interpreting Attention Layer Outputs with Sparse Autoencoders
por: Kissane, Connor, et al.
Publicado: (2024)
por: Kissane, Connor, et al.
Publicado: (2024)
Low-Rank Adapting Models for Sparse Autoencoders
por: Chen, Matthew, et al.
Publicado: (2025)
por: Chen, Matthew, et al.
Publicado: (2025)
Route Sparse Autoencoder to Interpret Large Language Models
por: Shi, Wei, et al.
Publicado: (2025)
por: Shi, Wei, et al.
Publicado: (2025)
Step-Level Sparse Autoencoder for Reasoning Process Interpretation
por: Yang, Xuan, et al.
Publicado: (2026)
por: Yang, Xuan, et al.
Publicado: (2026)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
por: Makelov, Aleksandar, et al.
Publicado: (2024)
por: Makelov, Aleksandar, et al.
Publicado: (2024)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
por: Cho, Hakaze, et al.
Publicado: (2025)
por: Cho, Hakaze, et al.
Publicado: (2025)
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
por: Roy, Dip, et al.
Publicado: (2025)
por: Roy, Dip, et al.
Publicado: (2025)
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
por: Zheng, Carolina, et al.
Publicado: (2025)
por: Zheng, Carolina, et al.
Publicado: (2025)
Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
por: Ye, Mengyu, et al.
Publicado: (2025)
por: Ye, Mengyu, et al.
Publicado: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
por: Marks, Luke, et al.
Publicado: (2024)
por: Marks, Luke, et al.
Publicado: (2024)
Sketching the Heat Kernel: Using Gaussian Processes to Embed Data
por: Gilbert, Anna C., et al.
Publicado: (2024)
por: Gilbert, Anna C., et al.
Publicado: (2024)
Interpretable Company Similarity with Sparse Autoencoders
por: Molinari, Marco, et al.
Publicado: (2024)
por: Molinari, Marco, et al.
Publicado: (2024)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
por: Kurochkin, Vadim, et al.
Publicado: (2025)
por: Kurochkin, Vadim, et al.
Publicado: (2025)
Efficient Dictionary Learning with Switch Sparse Autoencoders
por: Mudide, Anish, et al.
Publicado: (2024)
por: Mudide, Anish, et al.
Publicado: (2024)
From superposition to sparse codes: interpretable representations in neural networks
por: Klindt, David, et al.
Publicado: (2025)
por: Klindt, David, et al.
Publicado: (2025)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
por: Kantamneni, Subhash, et al.
Publicado: (2025)
por: Kantamneni, Subhash, et al.
Publicado: (2025)
Measuring Sharpness in Grokking
por: Miller, Jack, et al.
Publicado: (2024)
por: Miller, Jack, et al.
Publicado: (2024)
Interpreting CLIP with Hierarchical Sparse Autoencoders
por: Zaigrajew, Vladimir, et al.
Publicado: (2025)
por: Zaigrajew, Vladimir, et al.
Publicado: (2025)
Interpreting and Steering Protein Language Models through Sparse Autoencoders
por: Garcia, Edith Natalia Villegas, et al.
Publicado: (2025)
por: Garcia, Edith Natalia Villegas, et al.
Publicado: (2025)
Modelling the Doughnut of social and planetary boundaries with frugal machine learning
por: Vrizzi, Stefano, et al.
Publicado: (2025)
por: Vrizzi, Stefano, et al.
Publicado: (2025)
Interpreting CFD Surrogates through Sparse Autoencoders
por: Hu, Yeping, et al.
Publicado: (2025)
por: Hu, Yeping, et al.
Publicado: (2025)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
por: Kulkarni, Akshay, et al.
Publicado: (2025)
por: Kulkarni, Akshay, et al.
Publicado: (2025)
XNNTab -- Interpretable Neural Networks for Tabular Data using Sparse Autoencoders
por: Elhadri, Khawla, et al.
Publicado: (2025)
por: Elhadri, Khawla, et al.
Publicado: (2025)
Towards Interpretable Protein Structure Prediction with Sparse Autoencoders
por: Parsan, Nithin, et al.
Publicado: (2025)
por: Parsan, Nithin, et al.
Publicado: (2025)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
por: Zhu, Xudong, et al.
Publicado: (2025)
por: Zhu, Xudong, et al.
Publicado: (2025)
Ejemplares similares
-
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
por: O'Neill, Charles, et al.
Publicado: (2025) -
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
por: O'Neill, Charles, et al.
Publicado: (2024) -
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
por: O'Neill, Charles, et al.
Publicado: (2024) -
Disentangling Dense Embeddings with Sparse Autoencoders
por: O'Neill, Charles, et al.
Publicado: (2024) -
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
por: O'Neill, Charles
Publicado: (2025)