SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | He, Zirui, Jin, Mingyu, Shen, Bo, Payani, Ali, Zhang, Yongfeng, Du, Mengnan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
by: He, Zirui, et al.
Published: (2025)
by: He, Zirui, et al.
Published: (2025)
SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models
by: Han, Jiaojiao, et al.
Published: (2025)
by: Han, Jiaojiao, et al.
Published: (2025)
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection
by: Zhang, Huopu, et al.
Published: (2025)
by: Zhang, Huopu, et al.
Published: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation
by: Zhao, Haiyan, et al.
Published: (2026)
by: Zhao, Haiyan, et al.
Published: (2026)
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
by: He, Zirui, et al.
Published: (2025)
by: He, Zirui, et al.
Published: (2025)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
Knowledge Graph Large Language Model (KG-LLM) for Link Prediction
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
FinAnchor: Aligned Multi-Model Representations for Financial Prediction
by: He, Zirui, et al.
Published: (2026)
by: He, Zirui, et al.
Published: (2026)
Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
by: Zhao, Haiyan, et al.
Published: (2024)
by: Zhao, Haiyan, et al.
Published: (2024)
Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding
by: Jin, Mingyu, et al.
Published: (2025)
by: Jin, Mingyu, et al.
Published: (2025)
Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
The Impact of Reasoning Step Length on Large Language Models
by: Jin, Mingyu, et al.
Published: (2024)
by: Jin, Mingyu, et al.
Published: (2024)
Time Series Forecasting with LLMs: Understanding and Enhancing Model Capabilities
by: Tang, Hua, et al.
Published: (2024)
by: Tang, Hua, et al.
Published: (2024)
Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
The Cylindrical Representation Hypothesis for Language Model Steering
by: Gao, Lang, et al.
Published: (2026)
by: Gao, Lang, et al.
Published: (2026)
AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control
by: Li, Ruosen, et al.
Published: (2025)
by: Li, Ruosen, et al.
Published: (2025)
Towards Uncovering How Large Language Model Works: An Explainability Perspective
by: Zhao, Haiyan, et al.
Published: (2024)
by: Zhao, Haiyan, et al.
Published: (2024)
Reliable Control-Point Selection for Steering Reasoning in Large Language Models
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
Steering off Course: Reliability Challenges in Steering Language Models
by: Da Silva, Patrick Queiroz, et al.
Published: (2025)
by: Da Silva, Patrick Queiroz, et al.
Published: (2025)
LawLLM: Law Large Language Model for the US Legal System
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
Interpretable LLM Guardrails via Sparse Representation Steering
by: He, Zeqing, et al.
Published: (2025)
by: He, Zeqing, et al.
Published: (2025)
AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model
by: Xiong, Siheng, et al.
Published: (2024)
by: Xiong, Siheng, et al.
Published: (2024)
Large Language Models Can Learn Temporal Reasoning
by: Xiong, Siheng, et al.
Published: (2024)
by: Xiong, Siheng, et al.
Published: (2024)
Improved Representation Steering for Language Models
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
Health-LLM: Personalized Retrieval-Augmented Disease Prediction System
by: Yu, Qinkai, et al.
Published: (2024)
by: Yu, Qinkai, et al.
Published: (2024)
HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders
by: Chen, Boshui, et al.
Published: (2026)
by: Chen, Boshui, et al.
Published: (2026)
Improving LLM Reasoning through Interpretable Role-Playing Steering
by: Wang, Anyi, et al.
Published: (2025)
by: Wang, Anyi, et al.
Published: (2025)
Model Editing as a Double-Edged Sword: Steering Agent Ethical Behavior Toward Beneficence or Harm
by: Huang, Baixiang, et al.
Published: (2025)
by: Huang, Baixiang, et al.
Published: (2025)
Exploring Concept Depth: How Large Language Models Acquire Knowledge and Concept at Different Layers?
by: Jin, Mingyu, et al.
Published: (2024)
by: Jin, Mingyu, et al.
Published: (2024)
AlignSAE: Concept-Aligned Sparse Autoencoders
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groups
by: Ghilardi, Davide, et al.
Published: (2024)
by: Ghilardi, Davide, et al.
Published: (2024)
Disentangling Memory and Reasoning Ability in Large Language Models
by: Jin, Mingyu, et al.
Published: (2024)
by: Jin, Mingyu, et al.
Published: (2024)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
by: Chou, Cheng-Ting, et al.
Published: (2025)
by: Chou, Cheng-Ting, et al.
Published: (2025)
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
by: Wedgwood, James, et al.
Published: (2026)
by: Wedgwood, James, et al.
Published: (2026)
Enhancing Long Chain-of-Thought Reasoning through Multi-Path Plan Aggregation
by: Xiong, Siheng, et al.
Published: (2025)
by: Xiong, Siheng, et al.
Published: (2025)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
Similar Items
-
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
by: He, Zirui, et al.
Published: (2025) -
SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models
by: Han, Jiaojiao, et al.
Published: (2025) -
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection
by: Zhang, Huopu, et al.
Published: (2025) -
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
by: Zhao, Haiyan, et al.
Published: (2025) -
Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation
by: Zhao, Haiyan, et al.
Published: (2026)