Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | O'Neill, Charles, Jayasekara, Mudith, Kirkby, Max |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
Disentangling Dense Embeddings with Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
von: O'Neill, Charles
Veröffentlicht: (2025)
von: O'Neill, Charles
Veröffentlicht: (2025)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
von: Tolooshams, Bahareh, et al.
Veröffentlicht: (2025)
von: Tolooshams, Bahareh, et al.
Veröffentlicht: (2025)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
von: Erdogan, Ege, et al.
Veröffentlicht: (2025)
von: Erdogan, Ege, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
Type 2 Tobit Sample Selection Models with Bayesian Additive Regression Trees
von: O'Neill, Eoghan
Veröffentlicht: (2025)
von: O'Neill, Eoghan
Veröffentlicht: (2025)
Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
von: Lehn-Schiøler, William, et al.
Veröffentlicht: (2026)
von: Lehn-Schiøler, William, et al.
Veröffentlicht: (2026)
Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity
von: Miller, Jack, et al.
Veröffentlicht: (2023)
von: Miller, Jack, et al.
Veröffentlicht: (2023)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
von: Wang, Xu, et al.
Veröffentlicht: (2026)
von: Wang, Xu, et al.
Veröffentlicht: (2026)
Transcoders Beat Sparse Autoencoders for Interpretability
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
von: Engels, Joshua, et al.
Veröffentlicht: (2024)
von: Engels, Joshua, et al.
Veröffentlicht: (2024)
Interpretable Reward Model via Sparse Autoencoder
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
Interpreting Attention Layer Outputs with Sparse Autoencoders
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
Low-Rank Adapting Models for Sparse Autoencoders
von: Chen, Matthew, et al.
Veröffentlicht: (2025)
von: Chen, Matthew, et al.
Veröffentlicht: (2025)
Route Sparse Autoencoder to Interpret Large Language Models
von: Shi, Wei, et al.
Veröffentlicht: (2025)
von: Shi, Wei, et al.
Veröffentlicht: (2025)
Step-Level Sparse Autoencoder for Reasoning Process Interpretation
von: Yang, Xuan, et al.
Veröffentlicht: (2026)
von: Yang, Xuan, et al.
Veröffentlicht: (2026)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
von: Roy, Dip, et al.
Veröffentlicht: (2025)
von: Roy, Dip, et al.
Veröffentlicht: (2025)
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
von: Zheng, Carolina, et al.
Veröffentlicht: (2025)
von: Zheng, Carolina, et al.
Veröffentlicht: (2025)
Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
von: Ye, Mengyu, et al.
Veröffentlicht: (2025)
von: Ye, Mengyu, et al.
Veröffentlicht: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
von: Marks, Luke, et al.
Veröffentlicht: (2024)
von: Marks, Luke, et al.
Veröffentlicht: (2024)
Sketching the Heat Kernel: Using Gaussian Processes to Embed Data
von: Gilbert, Anna C., et al.
Veröffentlicht: (2024)
von: Gilbert, Anna C., et al.
Veröffentlicht: (2024)
Interpretable Company Similarity with Sparse Autoencoders
von: Molinari, Marco, et al.
Veröffentlicht: (2024)
von: Molinari, Marco, et al.
Veröffentlicht: (2024)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
von: Kurochkin, Vadim, et al.
Veröffentlicht: (2025)
von: Kurochkin, Vadim, et al.
Veröffentlicht: (2025)
Efficient Dictionary Learning with Switch Sparse Autoencoders
von: Mudide, Anish, et al.
Veröffentlicht: (2024)
von: Mudide, Anish, et al.
Veröffentlicht: (2024)
From superposition to sparse codes: interpretable representations in neural networks
von: Klindt, David, et al.
Veröffentlicht: (2025)
von: Klindt, David, et al.
Veröffentlicht: (2025)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Measuring Sharpness in Grokking
von: Miller, Jack, et al.
Veröffentlicht: (2024)
von: Miller, Jack, et al.
Veröffentlicht: (2024)
Interpreting CLIP with Hierarchical Sparse Autoencoders
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
Interpreting and Steering Protein Language Models through Sparse Autoencoders
von: Garcia, Edith Natalia Villegas, et al.
Veröffentlicht: (2025)
von: Garcia, Edith Natalia Villegas, et al.
Veröffentlicht: (2025)
Modelling the Doughnut of social and planetary boundaries with frugal machine learning
von: Vrizzi, Stefano, et al.
Veröffentlicht: (2025)
von: Vrizzi, Stefano, et al.
Veröffentlicht: (2025)
Interpreting CFD Surrogates through Sparse Autoencoders
von: Hu, Yeping, et al.
Veröffentlicht: (2025)
von: Hu, Yeping, et al.
Veröffentlicht: (2025)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
von: Kulkarni, Akshay, et al.
Veröffentlicht: (2025)
von: Kulkarni, Akshay, et al.
Veröffentlicht: (2025)
XNNTab -- Interpretable Neural Networks for Tabular Data using Sparse Autoencoders
von: Elhadri, Khawla, et al.
Veröffentlicht: (2025)
von: Elhadri, Khawla, et al.
Veröffentlicht: (2025)
Towards Interpretable Protein Structure Prediction with Sparse Autoencoders
von: Parsan, Nithin, et al.
Veröffentlicht: (2025)
von: Parsan, Nithin, et al.
Veröffentlicht: (2025)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
von: O'Neill, Charles, et al.
Veröffentlicht: (2025) -
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
von: O'Neill, Charles, et al.
Veröffentlicht: (2024) -
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2024) -
Disentangling Dense Embeddings with Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2024) -
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
von: O'Neill, Charles
Veröffentlicht: (2025)