Step-Level Sparse Autoencoder for Reasoning Process Interpretation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Xuan, Liu, Jiayu, Lai, Yuhang, Xu, Hao, Huang, Zhenya, Miao, Ning
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910278326681600
author Yang, Xuan
Liu, Jiayu
Lai, Yuhang
Xu, Hao
Huang, Zhenya
Miao, Ning
author_facet Yang, Xuan
Liu, Jiayu
Lai, Yuhang
Xu, Hao
Huang, Zhenya
Miao, Ning
contents Large Language Models (LLMs) have achieved strong complex reasoning capabilities through Chain-of-Thought (CoT) reasoning. However, their reasoning patterns remain too complicated to analyze. While Sparse Autoencoders (SAEs) have emerged as a powerful tool for interpretability, existing approaches predominantly operate at the token level, creating a granularity mismatch when capturing more critical step-level information, such as reasoning direction and semantic transitions. In this work, we propose step-level sparse autoencoder (SSAE), which serves as an analytical tool to disentangle different aspects of LLMs' reasoning steps into sparse features. Specifically, by precisely controlling the sparsity of a step feature conditioned on its context, we form an information bottleneck in step reconstruction, which splits incremental information from background information and disentangles it into several sparsely activated dimensions. Experiments on multiple base models and reasoning tasks show the effectiveness of the extracted features. By linear probing, we can easily predict surface-level information, such as generation length and first token distribution, as well as more complicated properties, such as the correctness and logicality of the step. These observations indicate that LLMs should already at least partly know about these properties during generation, which provides the foundation for the self-verification ability of LLMs. Our code is available at https://github.com/Miaow-Lab/SSAE.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03031
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Step-Level Sparse Autoencoder for Reasoning Process Interpretation
Yang, Xuan
Liu, Jiayu
Lai, Yuhang
Xu, Hao
Huang, Zhenya
Miao, Ning
Machine Learning
Large Language Models (LLMs) have achieved strong complex reasoning capabilities through Chain-of-Thought (CoT) reasoning. However, their reasoning patterns remain too complicated to analyze. While Sparse Autoencoders (SAEs) have emerged as a powerful tool for interpretability, existing approaches predominantly operate at the token level, creating a granularity mismatch when capturing more critical step-level information, such as reasoning direction and semantic transitions. In this work, we propose step-level sparse autoencoder (SSAE), which serves as an analytical tool to disentangle different aspects of LLMs' reasoning steps into sparse features. Specifically, by precisely controlling the sparsity of a step feature conditioned on its context, we form an information bottleneck in step reconstruction, which splits incremental information from background information and disentangles it into several sparsely activated dimensions. Experiments on multiple base models and reasoning tasks show the effectiveness of the extracted features. By linear probing, we can easily predict surface-level information, such as generation length and first token distribution, as well as more complicated properties, such as the correctness and logicality of the step. These observations indicate that LLMs should already at least partly know about these properties during generation, which provides the foundation for the self-verification ability of LLMs. Our code is available at https://github.com/Miaow-Lab/SSAE.
title Step-Level Sparse Autoencoder for Reasoning Process Interpretation
topic Machine Learning
url https://arxiv.org/abs/2603.03031