Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Renliang, Cheng, Wei, Li, Dawei, Chen, Haifeng, Wang, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908587617419264
author Sun, Renliang
Cheng, Wei
Li, Dawei
Chen, Haifeng
Wang, Wei
author_facet Sun, Renliang
Cheng, Wei
Li, Dawei
Chen, Haifeng
Wang, Wei
contents Chain-of-Thought (CoT) reasoning has driven recent gains of large language models (LLMs) on reasoning-intensive tasks by externalizing intermediate steps. However, excessive or redundant reasoning -- so-called overthinking -- can increase inference costs and lead LLMs toward incorrect conclusions. In this paper, we present REFRAIN ($\underline{REF}$lective-$\underline{R}$edundancy for $\underline{A}$daptive $\underline{IN}$ference), a training-free framework that adaptively determines when to stop reasoning to mitigate overthinking. REFRAIN integrates a two-stage stop discriminator to identify reflective yet redundant reasoning and a sliding-window Upper Confidence Bound (SW-UCB) multi-armed bandit controller to dynamically adjust stopping thresholds according to problem difficulty without supervision or fine-tuning. Across four representative benchmarks and two model families, REFRAIN reduces token usage by 20-55% while maintaining or improving accuracy compared to standard CoT prompting. Extensive ablation and robustness analyses demonstrate its stability across models, scorers, and prompt variations. In summary, our findings highlight when-to-stop as a new and practical axis of test-time scaling -- enabling models to reason not just more, but just enough.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10103
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
Sun, Renliang
Cheng, Wei
Li, Dawei
Chen, Haifeng
Wang, Wei
Computation and Language
Chain-of-Thought (CoT) reasoning has driven recent gains of large language models (LLMs) on reasoning-intensive tasks by externalizing intermediate steps. However, excessive or redundant reasoning -- so-called overthinking -- can increase inference costs and lead LLMs toward incorrect conclusions. In this paper, we present REFRAIN ($\underline{REF}$lective-$\underline{R}$edundancy for $\underline{A}$daptive $\underline{IN}$ference), a training-free framework that adaptively determines when to stop reasoning to mitigate overthinking. REFRAIN integrates a two-stage stop discriminator to identify reflective yet redundant reasoning and a sliding-window Upper Confidence Bound (SW-UCB) multi-armed bandit controller to dynamically adjust stopping thresholds according to problem difficulty without supervision or fine-tuning. Across four representative benchmarks and two model families, REFRAIN reduces token usage by 20-55% while maintaining or improving accuracy compared to standard CoT prompting. Extensive ablation and robustness analyses demonstrate its stability across models, scorers, and prompt variations. In summary, our findings highlight when-to-stop as a new and practical axis of test-time scaling -- enabling models to reason not just more, but just enough.
title Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
topic Computation and Language
url https://arxiv.org/abs/2510.10103