Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhenyu, Zhang, Shujian, Lambert, John, Zhou, Wenxuan, Wang, Zhangyang, Chen, Mingqing, Hard, Andrew, Mathews, Rajiv, Wang, Lun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917175513579520
author Zhang, Zhenyu
Zhang, Shujian
Lambert, John
Zhou, Wenxuan
Wang, Zhangyang
Chen, Mingqing
Hard, Andrew
Mathews, Rajiv
Wang, Lun
author_facet Zhang, Zhenyu
Zhang, Shujian
Lambert, John
Zhou, Wenxuan
Wang, Zhangyang
Chen, Mingqing
Hard, Andrew
Mathews, Rajiv
Wang, Lun
contents Despite the growing reasoning capabilities of recent large language models (LLMs), their internal mechanisms during the reasoning process remain underexplored. Prior approaches often rely on human-defined concepts (e.g., overthinking, reflection) at the word level to analyze reasoning in a supervised manner. However, such methods are limited, as it is infeasible to capture the full spectrum of potential reasoning behaviors, many of which are difficult to define in token space. In this work, we propose an unsupervised framework (namely, RISE: Reasoning behavior Interpretability via Sparse auto-Encoder) for discovering reasoning vectors, which we define as directions in the activation space that encode distinct reasoning behaviors. By segmenting chain-of-thought traces into sentence-level 'steps' and training sparse auto-encoders (SAEs) on step-level activations, we uncover disentangled features corresponding to interpretable behaviors such as reflection and backtracking. Visualization and clustering analyses show that these behaviors occupy separable regions in the decoder column space. Moreover, targeted interventions on SAE-derived vectors can controllably amplify or suppress specific reasoning behaviors, altering inference trajectories without retraining. Beyond behavior-specific disentanglement, SAEs capture structural properties such as response length, revealing clusters of long versus short reasoning traces. More interestingly, SAEs enable the discovery of novel behaviors beyond human supervision. We demonstrate the ability to control response confidence by identifying confidence-related vectors in the SAE decoder space. These findings underscore the potential of unsupervised latent discovery for both interpreting and controllably steering reasoning in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23988
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
Zhang, Zhenyu
Zhang, Shujian
Lambert, John
Zhou, Wenxuan
Wang, Zhangyang
Chen, Mingqing
Hard, Andrew
Mathews, Rajiv
Wang, Lun
Computation and Language
Artificial Intelligence
Machine Learning
Despite the growing reasoning capabilities of recent large language models (LLMs), their internal mechanisms during the reasoning process remain underexplored. Prior approaches often rely on human-defined concepts (e.g., overthinking, reflection) at the word level to analyze reasoning in a supervised manner. However, such methods are limited, as it is infeasible to capture the full spectrum of potential reasoning behaviors, many of which are difficult to define in token space. In this work, we propose an unsupervised framework (namely, RISE: Reasoning behavior Interpretability via Sparse auto-Encoder) for discovering reasoning vectors, which we define as directions in the activation space that encode distinct reasoning behaviors. By segmenting chain-of-thought traces into sentence-level 'steps' and training sparse auto-encoders (SAEs) on step-level activations, we uncover disentangled features corresponding to interpretable behaviors such as reflection and backtracking. Visualization and clustering analyses show that these behaviors occupy separable regions in the decoder column space. Moreover, targeted interventions on SAE-derived vectors can controllably amplify or suppress specific reasoning behaviors, altering inference trajectories without retraining. Beyond behavior-specific disentanglement, SAEs capture structural properties such as response length, revealing clusters of long versus short reasoning traces. More interestingly, SAEs enable the discovery of novel behaviors beyond human supervision. We demonstrate the ability to control response confidence by identifying confidence-related vectors in the SAE decoder space. These findings underscore the potential of unsupervised latent discovery for both interpreting and controllably steering reasoning in LLMs.
title Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.23988