Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Haiyan, Wu, Xuansheng, Yang, Fan, Shen, Bo, Liu, Ninghao, Du, Mengnan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912510241669120
author Zhao, Haiyan
Wu, Xuansheng
Yang, Fan
Shen, Bo
Liu, Ninghao
Du, Mengnan
author_facet Zhao, Haiyan
Wu, Xuansheng
Yang, Fan
Shen, Bo
Liu, Ninghao
Du, Mengnan
contents Linear concept vectors effectively steer LLMs, but existing methods suffer from noisy features in diverse datasets that undermine steering robustness. We propose Sparse Autoencoder-Denoised Concept Vectors (SDCV), which selectively keep the most discriminative SAE latents while reconstructing hidden representations. Our key insight is that concept-relevant signals can be explicitly separated from dataset noise by scaling up activations of top-k latents that best differentiate positive and negative samples. Applied to linear probing and difference-in-mean, SDCV consistently improves steering success rates by 4-16\% across six challenging concepts, while maintaining topic relevance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15038
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
Zhao, Haiyan
Wu, Xuansheng
Yang, Fan
Shen, Bo
Liu, Ninghao
Du, Mengnan
Computation and Language
Artificial Intelligence
Linear concept vectors effectively steer LLMs, but existing methods suffer from noisy features in diverse datasets that undermine steering robustness. We propose Sparse Autoencoder-Denoised Concept Vectors (SDCV), which selectively keep the most discriminative SAE latents while reconstructing hidden representations. Our key insight is that concept-relevant signals can be explicitly separated from dataset noise by scaling up activations of top-k latents that best differentiate positive and negative samples. Applied to linear probing and difference-in-mean, SDCV consistently improves steering success rates by 4-16\% across six challenging concepts, while maintaining topic relevance.
title Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.15038