TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nagle, Alliot, Saydaliev, Jakhongir, Garbaya, Dhia, Gastpar, Michael, Makkuva, Ashok Vardhan, Kim, Hyeji
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917492961574912
author Nagle, Alliot
Saydaliev, Jakhongir
Garbaya, Dhia
Gastpar, Michael
Makkuva, Ashok Vardhan
Kim, Hyeji
author_facet Nagle, Alliot
Saydaliev, Jakhongir
Garbaya, Dhia
Gastpar, Michael
Makkuva, Ashok Vardhan
Kim, Hyeji
contents Large Reasoning Models (LRMs) achieve impressive performance on complex reasoning tasks via Chain-of-Thought (CoT) reasoning, which enables them to generate intermediate thinking tokens before arriving at the final answer. However, LRMs often suffer from significant overthinking, spending excessive compute time even after the answer is generated early on. Prior work has identified the existence of an optimal reasoning length such that truncating reasoning at this point significantly shortens CoT outputs with virtually no change in performance. However, determining optimal CoT lengths for practical datasets is highly non-trivial as they are fully task and model-dependent. In this paper, we precisely address this and design Terminator, an early-exit strategy for LRMs at inference to mitigate overthinking. The central idea underpinning Terminator is that the first arrival of an LRM's final answer is often predictable, and we leverage these first answer positions to create a novel dataset of optimal reasoning lengths to train Terminator. Powered by this approach, Terminator achieves significant reductions in CoT lengths of 14%-55% on average across four challenging practical datasets: MATH-500, AIME 2025, HumanEval, and GPQA, while outperforming current state-of-the-art methods and reducing inference latency by more than 2x compared to the original LRM.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12529
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
Nagle, Alliot
Saydaliev, Jakhongir
Garbaya, Dhia
Gastpar, Michael
Makkuva, Ashok Vardhan
Kim, Hyeji
Machine Learning
Artificial Intelligence
Computation and Language
Large Reasoning Models (LRMs) achieve impressive performance on complex reasoning tasks via Chain-of-Thought (CoT) reasoning, which enables them to generate intermediate thinking tokens before arriving at the final answer. However, LRMs often suffer from significant overthinking, spending excessive compute time even after the answer is generated early on. Prior work has identified the existence of an optimal reasoning length such that truncating reasoning at this point significantly shortens CoT outputs with virtually no change in performance. However, determining optimal CoT lengths for practical datasets is highly non-trivial as they are fully task and model-dependent. In this paper, we precisely address this and design Terminator, an early-exit strategy for LRMs at inference to mitigate overthinking. The central idea underpinning Terminator is that the first arrival of an LRM's final answer is often predictable, and we leverage these first answer positions to create a novel dataset of optimal reasoning lengths to train Terminator. Powered by this approach, Terminator achieves significant reductions in CoT lengths of 14%-55% on average across four challenging practical datasets: MATH-500, AIME 2025, HumanEval, and GPQA, while outperforming current state-of-the-art methods and reducing inference latency by more than 2x compared to the original LRM.
title TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.12529