Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912510241669120 |
|---|---|
| author | Zhao, Haiyan Wu, Xuansheng Yang, Fan Shen, Bo Liu, Ninghao Du, Mengnan |
| author_facet | Zhao, Haiyan Wu, Xuansheng Yang, Fan Shen, Bo Liu, Ninghao Du, Mengnan |
| contents | Linear concept vectors effectively steer LLMs, but existing methods suffer from noisy features in diverse datasets that undermine steering robustness. We propose Sparse Autoencoder-Denoised Concept Vectors (SDCV), which selectively keep the most discriminative SAE latents while reconstructing hidden representations. Our key insight is that concept-relevant signals can be explicitly separated from dataset noise by scaling up activations of top-k latents that best differentiate positive and negative samples. Applied to linear probing and difference-in-mean, SDCV consistently improves steering success rates by 4-16\% across six challenging concepts, while maintaining topic relevance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_15038 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering Zhao, Haiyan Wu, Xuansheng Yang, Fan Shen, Bo Liu, Ninghao Du, Mengnan Computation and Language Artificial Intelligence Linear concept vectors effectively steer LLMs, but existing methods suffer from noisy features in diverse datasets that undermine steering robustness. We propose Sparse Autoencoder-Denoised Concept Vectors (SDCV), which selectively keep the most discriminative SAE latents while reconstructing hidden representations. Our key insight is that concept-relevant signals can be explicitly separated from dataset noise by scaling up activations of top-k latents that best differentiate positive and negative samples. Applied to linear probing and difference-in-mean, SDCV consistently improves steering success rates by 4-16\% across six challenging concepts, while maintaining topic relevance. |
| title | Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2505.15038 |