SpecExit: Accelerating Large Reasoning Model via Speculative Exit

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Rubing, Bai, Huajun, Liu, Song, Yu, Guanghua, Fan, Runzhi, Dang, Yanbin, Zhang, Jiejing, Liu, Kai, Zhu, Jianchen, Chen, Peng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915565125238784
author Yang, Rubing
Bai, Huajun
Liu, Song
Yu, Guanghua
Fan, Runzhi
Dang, Yanbin
Zhang, Jiejing
Liu, Kai
Zhu, Jianchen
Chen, Peng
author_facet Yang, Rubing
Bai, Huajun
Liu, Song
Yu, Guanghua
Fan, Runzhi
Dang, Yanbin
Zhang, Jiejing
Liu, Kai
Zhu, Jianchen
Chen, Peng
contents Despite their strong performance on reasoning tasks, large reasoning models (LRMs) often suffer from overthinking, producing unnecessarily long outputs and incurring high end-to-end latency, a significant limitation to their real-world deployment. To address overthinking, early-exit mechanisms have been proposed to terminate reasoning before typical completion, showing that this approach can effectively shorten generation length with minimal impact on accuracy. However, their reliance on probing mechanisms introduces a detection overhead that limits their end-to-end latency gains and compromises their generalizability across diverse problems. Inspired by the use of hidden states in speculative decoding, we propose SpecExit, a novel framework that predicts both future tokens and an early-exit signal directly from a lightweight draft model without probing overhead. Our method offers significant improvements, reducing average generation length by 66\% and achieving a 2.5x speedup in end-to-end latency compared to the speculative decoding baseline, without compromising accuracy. Our method leverages the inherent signals from hidden states to provide effective early-exit signals, suggesting broader use of hidden states for efficient reasoning. Our code is available at https://github.com/Tencent/AngelSlim.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24248
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpecExit: Accelerating Large Reasoning Model via Speculative Exit
Yang, Rubing
Bai, Huajun
Liu, Song
Yu, Guanghua
Fan, Runzhi
Dang, Yanbin
Zhang, Jiejing
Liu, Kai
Zhu, Jianchen
Chen, Peng
Artificial Intelligence
Computation and Language
Machine Learning
Despite their strong performance on reasoning tasks, large reasoning models (LRMs) often suffer from overthinking, producing unnecessarily long outputs and incurring high end-to-end latency, a significant limitation to their real-world deployment. To address overthinking, early-exit mechanisms have been proposed to terminate reasoning before typical completion, showing that this approach can effectively shorten generation length with minimal impact on accuracy. However, their reliance on probing mechanisms introduces a detection overhead that limits their end-to-end latency gains and compromises their generalizability across diverse problems. Inspired by the use of hidden states in speculative decoding, we propose SpecExit, a novel framework that predicts both future tokens and an early-exit signal directly from a lightweight draft model without probing overhead. Our method offers significant improvements, reducing average generation length by 66\% and achieving a 2.5x speedup in end-to-end latency compared to the speculative decoding baseline, without compromising accuracy. Our method leverages the inherent signals from hidden states to provide effective early-exit signals, suggesting broader use of hidden states for efficient reasoning. Our code is available at https://github.com/Tencent/AngelSlim.
title SpecExit: Accelerating Large Reasoning Model via Speculative Exit
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.24248