Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Anyi, Wu, Xuansheng, Shu, Dong, Ma, Yunpu, Liu, Ninghao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
Improving LLM Reasoning through Interpretable Role-Playing Steering
von: Wang, Anyi, et al.
Veröffentlicht: (2025)
von: Wang, Anyi, et al.
Veröffentlicht: (2025)
Graph-Regularized Sparse Autoencoders for LLM Safety Steering
von: Yeon, Jehyeok, et al.
Veröffentlicht: (2025)
von: Yeon, Jehyeok, et al.
Veröffentlicht: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
Retrieval-enhanced Knowledge Editing in Language Models for Multi-Hop Question Answering
von: Shi, Yucheng, et al.
Veröffentlicht: (2024)
von: Shi, Yucheng, et al.
Veröffentlicht: (2024)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
von: He, Zirui, et al.
Veröffentlicht: (2025)
von: He, Zirui, et al.
Veröffentlicht: (2025)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
von: Jørgensen, Mikkel Godsk, et al.
Veröffentlicht: (2026)
von: Jørgensen, Mikkel Godsk, et al.
Veröffentlicht: (2026)
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
Sparse Autoencoders, Again?
von: Lu, Yin, et al.
Veröffentlicht: (2025)
von: Lu, Yin, et al.
Veröffentlicht: (2025)
SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
von: Zimmer, Max, et al.
Veröffentlicht: (2025)
von: Zimmer, Max, et al.
Veröffentlicht: (2025)
Improving Sparse Autoencoder with Dynamic Attention
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
von: Hua, Zhenglin, et al.
Veröffentlicht: (2025)
von: Hua, Zhenglin, et al.
Veröffentlicht: (2025)
Are Sparse Autoencoder Benchmarks Reliable?
von: Chanin, David
Veröffentlicht: (2026)
von: Chanin, David
Veröffentlicht: (2026)
Dialz: A Python Toolkit for Steering Vectors
von: Siddique, Zara, et al.
Veröffentlicht: (2025)
von: Siddique, Zara, et al.
Veröffentlicht: (2025)
When LLM Reward Design Fails: Diagnostic-Driven Refinement for Sparse Structured RL
von: Wang, Youting, et al.
Veröffentlicht: (2026)
von: Wang, Youting, et al.
Veröffentlicht: (2026)
Steered LLM Activations are Non-Surjective
von: Mishra, Aayush, et al.
Veröffentlicht: (2026)
von: Mishra, Aayush, et al.
Veröffentlicht: (2026)
Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2025)
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2025)
BatchTopK Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification
von: Zhang, Lin, et al.
Veröffentlicht: (2025)
von: Zhang, Lin, et al.
Veröffentlicht: (2025)
EF-LLM: Energy Forecasting LLM with AI-assisted Automation, Enhanced Sparse Prediction, Hallucination Detection
von: Qiu, Zihang, et al.
Veröffentlicht: (2024)
von: Qiu, Zihang, et al.
Veröffentlicht: (2024)
On the Non-Identifiability of Steering Vectors in Large Language Models
von: Venkatesh, Sohan, et al.
Veröffentlicht: (2026)
von: Venkatesh, Sohan, et al.
Veröffentlicht: (2026)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
von: Ayonrinde, Kola
Veröffentlicht: (2024)
von: Ayonrinde, Kola
Veröffentlicht: (2024)
Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
von: Wu, Jialin, et al.
Veröffentlicht: (2026)
von: Wu, Jialin, et al.
Veröffentlicht: (2026)
BarrierSteer: LLM Safety via Learning Barrier Steering
von: Tran, Thanh Q., et al.
Veröffentlicht: (2026)
von: Tran, Thanh Q., et al.
Veröffentlicht: (2026)
Semantic Evolvement Enhanced Graph Autoencoder for Rumor Detection
von: Tao, Xiang, et al.
Veröffentlicht: (2024)
von: Tao, Xiang, et al.
Veröffentlicht: (2024)
Empirical Evaluation of Progressive Coding for Sparse Autoencoders
von: Peter, Hans, et al.
Veröffentlicht: (2025)
von: Peter, Hans, et al.
Veröffentlicht: (2025)
On the transferability of Sparse Autoencoders for interpreting compressed models
von: Gupte, Suchit, et al.
Veröffentlicht: (2025)
von: Gupte, Suchit, et al.
Veröffentlicht: (2025)
Data Whitening Improves Sparse Autoencoder Learning
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Diversity-driven Data Selection for Language Model Tuning through Sparse Autoencoder
von: Yang, Xianjun, et al.
Veröffentlicht: (2025)
von: Yang, Xianjun, et al.
Veröffentlicht: (2025)
Steering Large Language Model Activations in Sparse Spaces
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
von: Shu, Dong, et al.
Veröffentlicht: (2025) -
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
von: Shu, Dong, et al.
Veröffentlicht: (2025) -
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025) -
Improving LLM Reasoning through Interpretable Role-Playing Steering
von: Wang, Anyi, et al.
Veröffentlicht: (2025) -
Graph-Regularized Sparse Autoencoders for LLM Safety Steering
von: Yeon, Jehyeok, et al.
Veröffentlicht: (2025)