SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Chenzhi, Hu, Qinzhe, Xu, Yuhang, Chen, Junyi, Wang, Ruijie, Liu, Shengzhong, Li, Jianxin, Wu, Fan, Chen, Guihai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917550002012160
author Hu, Chenzhi
Hu, Qinzhe
Xu, Yuhang
Chen, Junyi
Wang, Ruijie
Liu, Shengzhong
Li, Jianxin
Wu, Fan
Chen, Guihai
author_facet Hu, Chenzhi
Hu, Qinzhe
Xu, Yuhang
Chen, Junyi
Wang, Ruijie
Liu, Shengzhong
Li, Jianxin
Wu, Fan
Chen, Guihai
contents Large reasoning models (LRMs) like OpenAI o1 and DeepSeek-R1 achieve high accuracy on complex tasks by adopting long chain-of-thought (CoT) reasoning paths. However, the inherent verbosity of these processes frequently results in redundancy and overthinking. To address this issue, existing works leverage Group Relative Policy Optimization (GRPO) to reduce LRM output length, but their static length reward design cannot dynamically adapt according to the relative problem difficulty and response length distribution, causing over-compression and compromised accuracy. Therefore, we propose SmartThinker, a novel GRPO-based efficient reasoning method with progressive CoT length calibration. SmartThinker makes a two-fold contribution: First, it dynamically estimates the optimal length with peak accuracy during training and guides overlong responses toward it to reduce response length while sustaining accuracy. Second, it dynamically modulates the length reward coefficient to avoid the unwarranted penalization of correct reasoning paths. Extensive experiment results show that SmartThinker achieves up to 52.5% average length compression with improved accuracy, and achieves up to 16.6% accuracy improvement on challenging benchmarks like AIME25. The source code can be found at https://github.com/SJTU-RTEAS/SmartThinker.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08000
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning
Hu, Chenzhi
Hu, Qinzhe
Xu, Yuhang
Chen, Junyi
Wang, Ruijie
Liu, Shengzhong
Li, Jianxin
Wu, Fan
Chen, Guihai
Computation and Language
Machine Learning
Large reasoning models (LRMs) like OpenAI o1 and DeepSeek-R1 achieve high accuracy on complex tasks by adopting long chain-of-thought (CoT) reasoning paths. However, the inherent verbosity of these processes frequently results in redundancy and overthinking. To address this issue, existing works leverage Group Relative Policy Optimization (GRPO) to reduce LRM output length, but their static length reward design cannot dynamically adapt according to the relative problem difficulty and response length distribution, causing over-compression and compromised accuracy. Therefore, we propose SmartThinker, a novel GRPO-based efficient reasoning method with progressive CoT length calibration. SmartThinker makes a two-fold contribution: First, it dynamically estimates the optimal length with peak accuracy during training and guides overlong responses toward it to reduce response length while sustaining accuracy. Second, it dynamically modulates the length reward coefficient to avoid the unwarranted penalization of correct reasoning paths. Extensive experiment results show that SmartThinker achieves up to 52.5% average length compression with improved accuracy, and achieves up to 16.6% accuracy improvement on challenging benchmarks like AIME25. The source code can be found at https://github.com/SJTU-RTEAS/SmartThinker.
title SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2603.08000