Saved in:
| Main Authors: | Yao, Yifei, Zhang, Hanrong, Du, Mengnan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.17320 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
by: He, Zirui, et al.
Published: (2025)
by: He, Zirui, et al.
Published: (2025)
Route Sparse Autoencoder to Interpret Large Language Models
by: Shi, Wei, et al.
Published: (2025)
by: Shi, Wei, et al.
Published: (2025)
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection
by: Zhang, Huopu, et al.
Published: (2025)
by: Zhang, Huopu, et al.
Published: (2025)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
Interpretable Reward Model via Sparse Autoencoder
by: Zhang, Shuyi, et al.
Published: (2025)
by: Zhang, Shuyi, et al.
Published: (2025)
Interpreting and Steering Protein Language Models through Sparse Autoencoders
by: Garcia, Edith Natalia Villegas, et al.
Published: (2025)
by: Garcia, Edith Natalia Villegas, et al.
Published: (2025)
Fine-Grained Interpretation of Political Opinions in Large Language Models
by: Hu, Jingyu, et al.
Published: (2025)
by: Hu, Jingyu, et al.
Published: (2025)
Sparse Autoencoders for Interpretable Medical Image Representation Learning
by: Wesp, Philipp, et al.
Published: (2026)
by: Wesp, Philipp, et al.
Published: (2026)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Transcoders Beat Sparse Autoencoders for Interpretability
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling
by: Miao, Yongliang, et al.
Published: (2026)
by: Miao, Yongliang, et al.
Published: (2026)
On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy
by: Cui, Jingyi, et al.
Published: (2025)
by: Cui, Jingyi, et al.
Published: (2025)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
by: Yamashita, Tomoya, et al.
Published: (2025)
by: Yamashita, Tomoya, et al.
Published: (2025)
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Interpreting Attention Layer Outputs with Sparse Autoencoders
by: Kissane, Connor, et al.
Published: (2024)
by: Kissane, Connor, et al.
Published: (2024)
Evaluating Sparse Autoencoders for Monosemantic Representation
by: Fereidouni, Moghis, et al.
Published: (2025)
by: Fereidouni, Moghis, et al.
Published: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
by: Wang, Xu, et al.
Published: (2026)
by: Wang, Xu, et al.
Published: (2026)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
by: Bhalla, Usha, et al.
Published: (2025)
by: Bhalla, Usha, et al.
Published: (2025)
Step-Level Sparse Autoencoder for Reasoning Process Interpretation
by: Yang, Xuan, et al.
Published: (2026)
by: Yang, Xuan, et al.
Published: (2026)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
by: Makelov, Aleksandar, et al.
Published: (2024)
by: Makelov, Aleksandar, et al.
Published: (2024)
Steering Language Model Refusal with Sparse Autoencoders
by: O'Brien, Kyle, et al.
Published: (2024)
by: O'Brien, Kyle, et al.
Published: (2024)
BatchTopK Sparse Autoencoders
by: Bussmann, Bart, et al.
Published: (2024)
by: Bussmann, Bart, et al.
Published: (2024)
SoftSAE: Dynamic Top-K Selection for Adaptive Sparse Autoencoders
by: Stępień, Jakub, et al.
Published: (2026)
by: Stępień, Jakub, et al.
Published: (2026)
Interpretable Company Similarity with Sparse Autoencoders
by: Molinari, Marco, et al.
Published: (2024)
by: Molinari, Marco, et al.
Published: (2024)
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
by: Yeung, Calvin, et al.
Published: (2026)
by: Yeung, Calvin, et al.
Published: (2026)
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025)
by: Zaigrajew, Vladimir, et al.
Published: (2025)
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
by: Cywiński, Bartosz, et al.
Published: (2025)
by: Cywiński, Bartosz, et al.
Published: (2025)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
by: Erdogan, Ege, et al.
Published: (2025)
by: Erdogan, Ege, et al.
Published: (2025)
Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
by: Ye, Mengyu, et al.
Published: (2025)
by: Ye, Mengyu, et al.
Published: (2025)
Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control
by: Klenitskiy, Anton, et al.
Published: (2025)
by: Klenitskiy, Anton, et al.
Published: (2025)
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
by: Thasarathan, Harrish, et al.
Published: (2025)
by: Thasarathan, Harrish, et al.
Published: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
by: Marks, Luke, et al.
Published: (2024)
by: Marks, Luke, et al.
Published: (2024)
Probing the Representational Power of Sparse Autoencoders in Vision Models
by: Olson, Matthew Lyle, et al.
Published: (2025)
by: Olson, Matthew Lyle, et al.
Published: (2025)
Interpreting CFD Surrogates through Sparse Autoencoders
by: Hu, Yeping, et al.
Published: (2025)
by: Hu, Yeping, et al.
Published: (2025)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
by: Tolooshams, Bahareh, et al.
Published: (2025)
by: Tolooshams, Bahareh, et al.
Published: (2025)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
by: Kulkarni, Akshay, et al.
Published: (2025)
by: Kulkarni, Akshay, et al.
Published: (2025)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025)
by: Kurochkin, Vadim, et al.
Published: (2025)
InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders
by: Simon, Elana, et al.
Published: (2024)
by: Simon, Elana, et al.
Published: (2024)
Aligning Large Language Models and Geometric Deep Models for Protein Representation
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
Similar Items
-
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
by: Shu, Dong, et al.
Published: (2025) -
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
by: He, Zirui, et al.
Published: (2025) -
Route Sparse Autoencoder to Interpret Large Language Models
by: Shi, Wei, et al.
Published: (2025) -
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection
by: Zhang, Huopu, et al.
Published: (2025) -
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
by: Shu, Dong, et al.
Published: (2025)