Controlling Thinking Speed in Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Zhengkai, Fu, Zhihang, Chen, Ze, Chen, Chao, Xie, Liang, Wang, Wenxiao, Cai, Deng, Wang, Zheng, Ye, Jieping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909877415182336
author Lin, Zhengkai
Fu, Zhihang
Chen, Ze
Chen, Chao
Xie, Liang
Wang, Wenxiao
Cai, Deng
Wang, Zheng
Ye, Jieping
author_facet Lin, Zhengkai
Fu, Zhihang
Chen, Ze
Chen, Chao
Xie, Liang
Wang, Wenxiao
Cai, Deng
Wang, Zheng
Ye, Jieping
contents Human cognition is theorized to operate in two modes: fast, intuitive System 1 thinking and slow, deliberate System 2 thinking. While current Large Reasoning Models (LRMs) excel at System 2 thinking, their inability to perform fast thinking leads to high computational overhead and latency. In this work, we enable LRMs to approximate human intelligence through dynamic thinking speed adjustment, optimizing accuracy-efficiency trade-offs. Our approach addresses two key questions: (1) how to control thinking speed in LRMs, and (2) when to adjust it for optimal performance. For the first question, we identify the steering vector that governs slow-fast thinking transitions in LRMs' representation space. Using this vector, we achieve the first representation editing-based test-time scaling effect, outperforming existing prompt-based scaling methods. For the second question, we apply real-time difficulty estimation to signal reasoning segments of varying complexity. Combining these techniques, we propose the first reasoning strategy that enables fast processing of easy steps and deeper analysis for complex reasoning. Without any training or additional cost, our plug-in module delivers an average +1.3% accuracy with -8.6% token usage across leading LRMs and advanced reasoning benchmarks. All of our algorithms are implemented based on vLLM and are expected to support broader applications and inspire future research.
format Preprint
id arxiv_https___arxiv_org_abs_2507_03704
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Controlling Thinking Speed in Reasoning Models
Lin, Zhengkai
Fu, Zhihang
Chen, Ze
Chen, Chao
Xie, Liang
Wang, Wenxiao
Cai, Deng
Wang, Zheng
Ye, Jieping
Computation and Language
Artificial Intelligence
Human cognition is theorized to operate in two modes: fast, intuitive System 1 thinking and slow, deliberate System 2 thinking. While current Large Reasoning Models (LRMs) excel at System 2 thinking, their inability to perform fast thinking leads to high computational overhead and latency. In this work, we enable LRMs to approximate human intelligence through dynamic thinking speed adjustment, optimizing accuracy-efficiency trade-offs. Our approach addresses two key questions: (1) how to control thinking speed in LRMs, and (2) when to adjust it for optimal performance. For the first question, we identify the steering vector that governs slow-fast thinking transitions in LRMs' representation space. Using this vector, we achieve the first representation editing-based test-time scaling effect, outperforming existing prompt-based scaling methods. For the second question, we apply real-time difficulty estimation to signal reasoning segments of varying complexity. Combining these techniques, we propose the first reasoning strategy that enables fast processing of easy steps and deeper analysis for complex reasoning. Without any training or additional cost, our plug-in module delivers an average +1.3% accuracy with -8.6% token usage across leading LRMs and advanced reasoning benchmarks. All of our algorithms are implemented based on vLLM and are expected to support broader applications and inspire future research.
title Controlling Thinking Speed in Reasoning Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.03704