Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tu, Songjun, Lin, Jiahao, Zhang, Qichao, Tian, Xiangyu, Li, Linjing, Lan, Xiangyuan, Zhao, Dongbin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916998965886976
author Tu, Songjun
Lin, Jiahao
Zhang, Qichao
Tian, Xiangyu
Li, Linjing
Lan, Xiangyuan
Zhao, Dongbin
author_facet Tu, Songjun
Lin, Jiahao
Zhang, Qichao
Tian, Xiangyu
Li, Linjing
Lan, Xiangyuan
Zhao, Dongbin
contents Large reasoning models (LRMs) are proficient at generating explicit, step-by-step reasoning sequences before producing final answers. However, such detailed reasoning can introduce substantial computational overhead and latency, particularly for simple problems. To address this over-thinking problem, we explore how to equip LRMs with adaptive thinking capabilities: enabling them to dynamically decide whether or not to engage in explicit reasoning based on problem complexity. Building on R1-style distilled models, we observe that inserting a simple ellipsis ("...") into the prompt can stochastically trigger either a thinking or no-thinking mode, revealing a latent controllability in the reasoning behavior. Leveraging this property, we propose AutoThink, a multi-stage reinforcement learning (RL) framework that progressively optimizes reasoning policies via stage-wise reward shaping. AutoThink learns to invoke explicit reasoning only when necessary, while defaulting to succinct responses for simpler tasks. Experiments on five mainstream mathematical benchmarks demonstrate that AutoThink achieves favorable accuracy-efficiency trade-offs compared to recent prompting and RL-based pruning methods. It can be seamlessly integrated into any R1-style model, including both distilled and further fine-tuned variants. Notably, AutoThink improves relative accuracy by 6.4 percent while reducing token usage by 52 percent on DeepSeek-R1-Distill-Qwen-1.5B, establishing a scalable and adaptive reasoning paradigm for LRMs. Project Page: https://github.com/ScienceOne-AI/AutoThink.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
Tu, Songjun
Lin, Jiahao
Zhang, Qichao
Tian, Xiangyu
Li, Linjing
Lan, Xiangyuan
Zhao, Dongbin
Computation and Language
Artificial Intelligence
68T50
I.2.7
Large reasoning models (LRMs) are proficient at generating explicit, step-by-step reasoning sequences before producing final answers. However, such detailed reasoning can introduce substantial computational overhead and latency, particularly for simple problems. To address this over-thinking problem, we explore how to equip LRMs with adaptive thinking capabilities: enabling them to dynamically decide whether or not to engage in explicit reasoning based on problem complexity. Building on R1-style distilled models, we observe that inserting a simple ellipsis ("...") into the prompt can stochastically trigger either a thinking or no-thinking mode, revealing a latent controllability in the reasoning behavior. Leveraging this property, we propose AutoThink, a multi-stage reinforcement learning (RL) framework that progressively optimizes reasoning policies via stage-wise reward shaping. AutoThink learns to invoke explicit reasoning only when necessary, while defaulting to succinct responses for simpler tasks. Experiments on five mainstream mathematical benchmarks demonstrate that AutoThink achieves favorable accuracy-efficiency trade-offs compared to recent prompting and RL-based pruning methods. It can be seamlessly integrated into any R1-style model, including both distilled and further fine-tuned variants. Notably, AutoThink improves relative accuracy by 6.4 percent while reducing token usage by 52 percent on DeepSeek-R1-Distill-Qwen-1.5B, establishing a scalable and adaptive reasoning paradigm for LRMs. Project Page: https://github.com/ScienceOne-AI/AutoThink.
title Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
topic Computation and Language
Artificial Intelligence
68T50
I.2.7
url https://arxiv.org/abs/2505.10832