CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cho, Seonglae, Wu, Zekun, Koshiyama, Adriano
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915974624575488
author Cho, Seonglae
Wu, Zekun
Koshiyama, Adriano
author_facet Cho, Seonglae
Wu, Zekun
Koshiyama, Adriano
contents Sparse Autoencoders (SAEs) can extract interpretable features from large language models (LLMs) without supervision. However, their effectiveness in downstream steering tasks is limited by the requirement for contrastive datasets or large activation storage. To address these limitations, we propose CorrSteer, which selects features by correlating sample correctness with SAE activations from generated tokens at inference time. This approach uses only inference-time activations to extract more relevant features, thereby reducing spurious correlations. It also obtains steering coefficients from average activations, automating the entire pipeline. Our method shows improved task performance on QA, bias mitigation, jailbreaking prevention, and reasoning benchmarks on Gemma-2 2B and LLaMA-3.1 8B, notably achieving a +3.3% improvement in MMLU performance with 4000 samples and a +27.2% improvement in HarmBench with only 108 samples. Selected features demonstrate semantically meaningful patterns aligned with each task's requirements, revealing the underlying capabilities that drive performance. Our work establishes correlation-based selection as an effective and scalable approach for automated SAE steering across language model applications.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12535
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
Cho, Seonglae
Wu, Zekun
Koshiyama, Adriano
Computation and Language
Artificial Intelligence
Machine Learning
I.2.7; I.2.6
Sparse Autoencoders (SAEs) can extract interpretable features from large language models (LLMs) without supervision. However, their effectiveness in downstream steering tasks is limited by the requirement for contrastive datasets or large activation storage. To address these limitations, we propose CorrSteer, which selects features by correlating sample correctness with SAE activations from generated tokens at inference time. This approach uses only inference-time activations to extract more relevant features, thereby reducing spurious correlations. It also obtains steering coefficients from average activations, automating the entire pipeline. Our method shows improved task performance on QA, bias mitigation, jailbreaking prevention, and reasoning benchmarks on Gemma-2 2B and LLaMA-3.1 8B, notably achieving a +3.3% improvement in MMLU performance with 4000 samples and a +27.2% improvement in HarmBench with only 108 samples. Selected features demonstrate semantically meaningful patterns aligned with each task's requirements, revealing the underlying capabilities that drive performance. Our work establishes correlation-based selection as an effective and scalable approach for automated SAE steering across language model applications.
title CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
topic Computation and Language
Artificial Intelligence
Machine Learning
I.2.7; I.2.6
url https://arxiv.org/abs/2508.12535