Route Sparse Autoencoder to Interpret Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Wei, Li, Sihang, Liang, Tao, Wan, Mingyang, Ma, Guojun, Wang, Xiang, He, Xiangnan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpretable Reward Model via Sparse Autoencoder
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
Towards 3D Molecule-Text Interpretation in Language Models
von: Li, Sihang, et al.
Veröffentlicht: (2024)
von: Li, Sihang, et al.
Veröffentlicht: (2024)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
von: Wang, Xu, et al.
Veröffentlicht: (2026)
von: Wang, Xu, et al.
Veröffentlicht: (2026)
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
von: Ma, George, et al.
Veröffentlicht: (2026)
von: Ma, George, et al.
Veröffentlicht: (2026)
Interpreting and Steering Protein Language Models through Sparse Autoencoders
von: Garcia, Edith Natalia Villegas, et al.
Veröffentlicht: (2025)
von: Garcia, Edith Natalia Villegas, et al.
Veröffentlicht: (2025)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
von: He, Zirui, et al.
Veröffentlicht: (2025)
von: He, Zirui, et al.
Veröffentlicht: (2025)
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
von: Weng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Weng, Jiaqi, et al.
Veröffentlicht: (2025)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations
von: Yao, Yifei, et al.
Veröffentlicht: (2025)
von: Yao, Yifei, et al.
Veröffentlicht: (2025)
Transcoders Beat Sparse Autoencoders for Interpretability
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
Interpreting Attention Layer Outputs with Sparse Autoencoders
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models
von: Demircan, Can, et al.
Veröffentlicht: (2024)
von: Demircan, Can, et al.
Veröffentlicht: (2024)
SAFER: Probing Safety in Reward Models with Sparse Autoencoder
von: Shi, Wei, et al.
Veröffentlicht: (2025)
von: Shi, Wei, et al.
Veröffentlicht: (2025)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
von: Kulkarni, Akshay, et al.
Veröffentlicht: (2025)
von: Kulkarni, Akshay, et al.
Veröffentlicht: (2025)
Q-Sparse: All Large Language Models can be Fully Sparsely-Activated
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
Steering Language Model Refusal with Sparse Autoencoders
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
Step-Level Sparse Autoencoder for Reasoning Process Interpretation
von: Yang, Xuan, et al.
Veröffentlicht: (2026)
von: Yang, Xuan, et al.
Veröffentlicht: (2026)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
von: Ma, Qingsen, et al.
Veröffentlicht: (2025)
von: Ma, Qingsen, et al.
Veröffentlicht: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
von: Bhalla, Usha, et al.
Veröffentlicht: (2025)
von: Bhalla, Usha, et al.
Veröffentlicht: (2025)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
von: Yamashita, Tomoya, et al.
Veröffentlicht: (2025)
von: Yamashita, Tomoya, et al.
Veröffentlicht: (2025)
Rethinking Graph Masked Autoencoders through Alignment and Uniformity
von: Wang, Liang, et al.
Veröffentlicht: (2024)
von: Wang, Liang, et al.
Veröffentlicht: (2024)
SparseEval: Efficient Evaluation of Large Language Models by Sparse Optimization
von: Zhang, Taolin, et al.
Veröffentlicht: (2026)
von: Zhang, Taolin, et al.
Veröffentlicht: (2026)
SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference
von: Ma, Hao, et al.
Veröffentlicht: (2026)
von: Ma, Hao, et al.
Veröffentlicht: (2026)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
von: Erdogan, Ege, et al.
Veröffentlicht: (2025)
von: Erdogan, Ege, et al.
Veröffentlicht: (2025)
Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
von: Ye, Mengyu, et al.
Veröffentlicht: (2025)
von: Ye, Mengyu, et al.
Veröffentlicht: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
von: Marks, Luke, et al.
Veröffentlicht: (2024)
von: Marks, Luke, et al.
Veröffentlicht: (2024)
In-context Autoencoder for Context Compression in a Large Language Model
von: Ge, Tao, et al.
Veröffentlicht: (2023)
von: Ge, Tao, et al.
Veröffentlicht: (2023)
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
von: Ding, Chenlu, et al.
Veröffentlicht: (2025)
von: Ding, Chenlu, et al.
Veröffentlicht: (2025)
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
Interpretable Company Similarity with Sparse Autoencoders
von: Molinari, Marco, et al.
Veröffentlicht: (2024)
von: Molinari, Marco, et al.
Veröffentlicht: (2024)
Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control
von: Klenitskiy, Anton, et al.
Veröffentlicht: (2025)
von: Klenitskiy, Anton, et al.
Veröffentlicht: (2025)
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
von: Thasarathan, Harrish, et al.
Veröffentlicht: (2025)
von: Thasarathan, Harrish, et al.
Veröffentlicht: (2025)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
von: Kurochkin, Vadim, et al.
Veröffentlicht: (2025)
von: Kurochkin, Vadim, et al.
Veröffentlicht: (2025)
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
von: Yeung, Calvin, et al.
Veröffentlicht: (2026)
von: Yeung, Calvin, et al.
Veröffentlicht: (2026)
Sparse Autoencoders, Again?
von: Lu, Yin, et al.
Veröffentlicht: (2025)
von: Lu, Yin, et al.
Veröffentlicht: (2025)
Interpreting CLIP with Hierarchical Sparse Autoencoders
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Interpretable Reward Model via Sparse Autoencoder
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025) -
Towards 3D Molecule-Text Interpretation in Language Models
von: Li, Sihang, et al.
Veröffentlicht: (2024) -
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
von: Wang, Xu, et al.
Veröffentlicht: (2026) -
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
von: Ma, George, et al.
Veröffentlicht: (2026) -
Interpreting and Steering Protein Language Models through Sparse Autoencoders
von: Garcia, Edith Natalia Villegas, et al.
Veröffentlicht: (2025)