Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Han, Jinyi, Huang, Ying, Liao, Ying, Jiang, Zishang, Lu, Xikun, Zhao, Haiquan, Wang, Xinyi, Zhou, Guanghao, Jiang, Sihang, Liang, Jiaqing, Zhou, Weikang, Sun, Zeye, Yu, Fei, Xiao, Yanghua
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914430951882752
author Han, Jinyi
Huang, Ying
Liao, Ying
Jiang, Zishang
Lu, Xikun
Zhao, Haiquan
Wang, Xinyi
Zhou, Guanghao
Jiang, Sihang
Liang, Jiaqing
Zhou, Weikang
Sun, Zeye
Yu, Fei
Xiao, Yanghua
author_facet Han, Jinyi
Huang, Ying
Liao, Ying
Jiang, Zishang
Lu, Xikun
Zhao, Haiquan
Wang, Xinyi
Zhou, Guanghao
Jiang, Sihang
Liang, Jiaqing
Zhou, Weikang
Sun, Zeye
Yu, Fei
Xiao, Yanghua
contents Large Reasoning Models (LRMs) have achieved impressive performance on challenging tasks, yet their deep reasoning often incurs substantial computational costs. To achieve efficient reasoning, existing reinforcement learning methods still struggle to construct short reasoning path during the rollout stage, limiting effective learning. Inspired by Evidence Accumulation Models, we find that LRMs have accumulated sufficient information early in reasoning, making further reasoning steps redundant. Based on this insight, we propose Just-Enough Thinking (JET), which trains models to proactively terminate unnecessary reasoning. JET performs trajectory truncation during rollout to expose the model to short, distributionally consistent reasoning paths. Besides, it uses a quality-controlled length reward to better encourage concise reasoning while maintaining correctness. Extensive experiments demonstrate that JET significantly improves reasoning efficiency without sacrificing accuracy. Especially, DeepSeek-Distill-Qwen-1.5B achieves a 4.6% accuracy gain while reducing output length by 46.3% on the Olympiad benchmark. Our code is available in the GitHub.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23392
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking
Han, Jinyi
Huang, Ying
Liao, Ying
Jiang, Zishang
Lu, Xikun
Zhao, Haiquan
Wang, Xinyi
Zhou, Guanghao
Jiang, Sihang
Liang, Jiaqing
Zhou, Weikang
Sun, Zeye
Yu, Fei
Xiao, Yanghua
Artificial Intelligence
Computation and Language
Large Reasoning Models (LRMs) have achieved impressive performance on challenging tasks, yet their deep reasoning often incurs substantial computational costs. To achieve efficient reasoning, existing reinforcement learning methods still struggle to construct short reasoning path during the rollout stage, limiting effective learning. Inspired by Evidence Accumulation Models, we find that LRMs have accumulated sufficient information early in reasoning, making further reasoning steps redundant. Based on this insight, we propose Just-Enough Thinking (JET), which trains models to proactively terminate unnecessary reasoning. JET performs trajectory truncation during rollout to expose the model to short, distributionally consistent reasoning paths. Besides, it uses a quality-controlled length reward to better encourage concise reasoning while maintaining correctness. Extensive experiments demonstrate that JET significantly improves reasoning efficiency without sacrificing accuracy. Especially, DeepSeek-Distill-Qwen-1.5B achieves a 4.6% accuracy gain while reducing output length by 46.3% on the Olympiad benchmark. Our code is available in the GitHub.
title Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.23392