Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zihao, Wang, Xu, Yang, Yuzhe, Yao, Ziyu, Xiong, Haoyi, Du, Mengnan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913938719899648
author Li, Zihao
Wang, Xu
Yang, Yuzhe
Yao, Ziyu
Xiong, Haoyi
Du, Mengnan
author_facet Li, Zihao
Wang, Xu
Yang, Yuzhe
Yao, Ziyu
Xiong, Haoyi
Du, Mengnan
contents Large Language Models (LLMs) demonstrate the ability to solve reasoning and mathematical problems using the Chain-of-Thought (CoT) technique. Expanding CoT length, as seen in models such as DeepSeek-R1, significantly enhances this reasoning for complex problems, but requires costly and high-quality long CoT data and fine-tuning. This work, inspired by the deep thinking paradigm of DeepSeek-R1, utilizes a steering technique to enhance the reasoning ability of an LLM without external datasets. Our method first employs Sparse Autoencoders (SAEs) to extract interpretable features from vanilla CoT. These features are then used to steer the LLM's internal states during generation. Recognizing that many LLMs do not have corresponding pre-trained SAEs, we further introduce a novel SAE-free steering algorithm, which directly computes steering directions from the residual activations of an LLM, obviating the need for an explicit SAE. Experimental results demonstrate that both our SAE-based and subsequent SAE-free steering algorithms significantly enhance the reasoning capabilities of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15634
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
Li, Zihao
Wang, Xu
Yang, Yuzhe
Yao, Ziyu
Xiong, Haoyi
Du, Mengnan
Computation and Language
Machine Learning
Large Language Models (LLMs) demonstrate the ability to solve reasoning and mathematical problems using the Chain-of-Thought (CoT) technique. Expanding CoT length, as seen in models such as DeepSeek-R1, significantly enhances this reasoning for complex problems, but requires costly and high-quality long CoT data and fine-tuning. This work, inspired by the deep thinking paradigm of DeepSeek-R1, utilizes a steering technique to enhance the reasoning ability of an LLM without external datasets. Our method first employs Sparse Autoencoders (SAEs) to extract interpretable features from vanilla CoT. These features are then used to steer the LLM's internal states during generation. Recognizing that many LLMs do not have corresponding pre-trained SAEs, we further introduce a novel SAE-free steering algorithm, which directly computes steering directions from the residual activations of an LLM, obviating the need for an explicit SAE. Experimental results demonstrate that both our SAE-based and subsequent SAE-free steering algorithms significantly enhance the reasoning capabilities of LLMs.
title Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.15634