Dynamic Early Exit in Reasoning Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Chenxu, Si, Qingyi, Duan, Yongjie, Zhu, Zheliang, Zhu, Chenyu, Li, Qiaowei, Chen, Minghui, Lin, Zheng, Wang, Weiping
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912610604023808
author Yang, Chenxu
Si, Qingyi
Duan, Yongjie
Zhu, Zheliang
Zhu, Chenyu
Li, Qiaowei
Chen, Minghui
Lin, Zheng
Wang, Weiping
author_facet Yang, Chenxu
Si, Qingyi
Duan, Yongjie
Zhu, Zheliang
Zhu, Chenyu
Li, Qiaowei
Chen, Minghui
Lin, Zheng
Wang, Weiping
contents Recent advances in large reasoning language models (LRLMs) rely on test-time scaling, which extends long chain-of-thought (CoT) generation to solve complex tasks. However, overthinking in long CoT not only slows down the efficiency of problem solving, but also risks accuracy loss due to the extremely detailed or redundant reasoning steps. We propose a simple yet effective method that allows LLMs to self-truncate CoT sequences by early exit during generation. Instead of relying on fixed heuristics, the proposed method monitors model behavior at potential reasoning transition points and dynamically terminates the next reasoning chain's generation when the model exhibits high confidence in a trial answer. Our method requires no additional training and can be seamlessly integrated into existing o1-like reasoning LLMs. Experiments on 10 reasoning benchmarks (e.g., GSM8K, MATH-500, AMC, GPQA, AIME and LiveCodeBench) show that the proposed method is consistently effective on 11 cutting-edge reasoning LLMs of varying series and sizes, reducing the length of CoT sequences by an average of 19.1% to 80.1% while improving accuracy by 0.3% to 5.0%.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15895
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Early Exit in Reasoning Models
Yang, Chenxu
Si, Qingyi
Duan, Yongjie
Zhu, Zheliang
Zhu, Chenyu
Li, Qiaowei
Chen, Minghui
Lin, Zheng
Wang, Weiping
Computation and Language
Artificial Intelligence
Recent advances in large reasoning language models (LRLMs) rely on test-time scaling, which extends long chain-of-thought (CoT) generation to solve complex tasks. However, overthinking in long CoT not only slows down the efficiency of problem solving, but also risks accuracy loss due to the extremely detailed or redundant reasoning steps. We propose a simple yet effective method that allows LLMs to self-truncate CoT sequences by early exit during generation. Instead of relying on fixed heuristics, the proposed method monitors model behavior at potential reasoning transition points and dynamically terminates the next reasoning chain's generation when the model exhibits high confidence in a trial answer. Our method requires no additional training and can be seamlessly integrated into existing o1-like reasoning LLMs. Experiments on 10 reasoning benchmarks (e.g., GSM8K, MATH-500, AMC, GPQA, AIME and LiveCodeBench) show that the proposed method is consistently effective on 11 cutting-edge reasoning LLMs of varying series and sizes, reducing the length of CoT sequences by an average of 19.1% to 80.1% while improving accuracy by 0.3% to 5.0%.
title Dynamic Early Exit in Reasoning Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.15895