A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Shu, Dong, Wu, Xuansheng, Zhao, Haiyan, Rai, Daking, Yao, Ziyu, Liu, Ninghao, Du, Mengnan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
Mechanistic Understanding of Language Models in Syntactic Code Completion
di: Miller, Samuel, et al.
Pubblicazione: (2025)
di: Miller, Samuel, et al.
Pubblicazione: (2025)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
di: Rai, Daking, et al.
Pubblicazione: (2024)
di: Rai, Daking, et al.
Pubblicazione: (2024)
Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
di: Rai, Daking, et al.
Pubblicazione: (2025)
di: Rai, Daking, et al.
Pubblicazione: (2025)
Data-driven Circuit Discovery for Interpretability of Language Models
di: Rai, Daking, et al.
Pubblicazione: (2026)
di: Rai, Daking, et al.
Pubblicazione: (2026)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
di: He, Zirui, et al.
Pubblicazione: (2025)
di: He, Zirui, et al.
Pubblicazione: (2025)
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
di: Wang, Anyi, et al.
Pubblicazione: (2025)
di: Wang, Anyi, et al.
Pubblicazione: (2025)
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
di: Lee, Gyeong-Geon, et al.
Pubblicazione: (2023)
di: Lee, Gyeong-Geon, et al.
Pubblicazione: (2023)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
di: Rai, Daking, et al.
Pubblicazione: (2024)
di: Rai, Daking, et al.
Pubblicazione: (2024)
AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
di: Wang, Yun, et al.
Pubblicazione: (2025)
di: Wang, Yun, et al.
Pubblicazione: (2025)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis
di: Li, Daoyang, et al.
Pubblicazione: (2024)
di: Li, Daoyang, et al.
Pubblicazione: (2024)
The Impact of Reasoning Step Length on Large Language Models
di: Jin, Mingyu, et al.
Pubblicazione: (2024)
di: Jin, Mingyu, et al.
Pubblicazione: (2024)
Could Small Language Models Serve as Recommenders? Towards Data-centric Cold-start Recommendations
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
di: Wu, Xuansheng, et al.
Pubblicazione: (2023)
Retrieval-enhanced Knowledge Editing in Language Models for Multi-Hop Question Answering
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
di: Wang, Yun, et al.
Pubblicazione: (2026)
di: Wang, Yun, et al.
Pubblicazione: (2026)
Fine-Grained Interpretation of Political Opinions in Large Language Models
di: Hu, Jingyu, et al.
Pubblicazione: (2025)
di: Hu, Jingyu, et al.
Pubblicazione: (2025)
A Survey on Fairness in Large Language Models
di: Li, Yingji, et al.
Pubblicazione: (2023)
di: Li, Yingji, et al.
Pubblicazione: (2023)
Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
di: Mamidanna, Siddarth, et al.
Pubblicazione: (2025)
di: Mamidanna, Siddarth, et al.
Pubblicazione: (2025)
Comparative Analysis of Demonstration Selection Algorithms for LLM In-Context Learning
di: Shu, Dong, et al.
Pubblicazione: (2024)
di: Shu, Dong, et al.
Pubblicazione: (2024)
Investigating CoT Monitorability in Large Reasoning Models
di: Yang, Shu, et al.
Pubblicazione: (2025)
di: Yang, Shu, et al.
Pubblicazione: (2025)
Improving LLM Reasoning through Interpretable Role-Playing Steering
di: Wang, Anyi, et al.
Pubblicazione: (2025)
di: Wang, Anyi, et al.
Pubblicazione: (2025)
Understanding the Effect of Algorithm Transparency of Model Explanations in Text-to-SQL Semantic Parsing
di: Rai, Daking, et al.
Pubblicazione: (2024)
di: Rai, Daking, et al.
Pubblicazione: (2024)
Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
di: Li, Zihao, et al.
Pubblicazione: (2024)
di: Li, Zihao, et al.
Pubblicazione: (2024)
A Survey on Large Language Models for Automated Planning
di: Aghzal, Mohamed, et al.
Pubblicazione: (2025)
di: Aghzal, Mohamed, et al.
Pubblicazione: (2025)
InFoBench: Evaluating Instruction Following Ability in Large Language Models
di: Qin, Yiwei, et al.
Pubblicazione: (2024)
di: Qin, Yiwei, et al.
Pubblicazione: (2024)
ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders
di: Liu, Xiangyu, et al.
Pubblicazione: (2025)
di: Liu, Xiangyu, et al.
Pubblicazione: (2025)
AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations
di: Yao, Yifei, et al.
Pubblicazione: (2025)
di: Yao, Yifei, et al.
Pubblicazione: (2025)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
di: Jing, Yi, et al.
Pubblicazione: (2026)
di: Jing, Yi, et al.
Pubblicazione: (2026)
Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
Concept-Centric Token Interpretation for Vector-Quantized Generative Models
di: Yang, Tianze, et al.
Pubblicazione: (2025)
di: Yang, Tianze, et al.
Pubblicazione: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
Towards Uncovering How Large Language Model Works: An Explainability Perspective
di: Zhao, Haiyan, et al.
Pubblicazione: (2024)
di: Zhao, Haiyan, et al.
Pubblicazione: (2024)
Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning
di: Yue, Murong, et al.
Pubblicazione: (2023)
di: Yue, Murong, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
di: Shu, Dong, et al.
Pubblicazione: (2025) -
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025) -
Mechanistic Understanding of Language Models in Syntactic Code Completion
di: Miller, Samuel, et al.
Pubblicazione: (2025) -
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
di: Wu, Xuansheng, et al.
Pubblicazione: (2025) -
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
di: Rai, Daking, et al.
Pubblicazione: (2024)