Beyond Token Length: Step Pruner for Efficient and Accurate Reasoning in Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wu, Canhui, Cao, Qiong, Li, Chang, Wang, Zhenfang, Xue, Chao, Fan, Yuwei, Xi, Wei, He, Xiaodong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911293482467328
author Wu, Canhui
Cao, Qiong
Li, Chang
Wang, Zhenfang
Xue, Chao
Fan, Yuwei
Xi, Wei
He, Xiaodong
author_facet Wu, Canhui
Cao, Qiong
Li, Chang
Wang, Zhenfang
Xue, Chao
Fan, Yuwei
Xi, Wei
He, Xiaodong
contents Large Reasoning Models (LRMs) demonstrate strong performance on complex tasks but often suffer from excessive verbosity, known as "overthinking." Existing solutions via reinforcement learning (RL) typically penalize generated tokens to promote conciseness. However, these methods encounter two challenges: responses with fewer tokens do not always correspond to fewer reasoning steps, and models may develop hacking behavior in later stages of training by discarding reasoning steps to minimize token usage. In this work, we introduce \textbf{Step Pruner (SP)}, an RL framework that steers LRMs toward more efficient reasoning by favoring compact reasoning steps. Our step-aware reward function prioritizes correctness while imposing penalties for redundant steps, and withholds rewards for incorrect responses to prevent the reinforcement of erroneous reasoning. Moreover, we propose a dynamic stopping mechanism: when the model's output no longer shortens, training is halted to prevent hacking behavior caused by the merging of steps. Extensive experiments across four reasoning benchmarks demonstrate that SP achieves state-of-the-art accuracy while significantly reducing response length. For instance, on AIME24, SP reduces token usage by \textbf{69.7\%}.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Token Length: Step Pruner for Efficient and Accurate Reasoning in Large Language Models
Wu, Canhui
Cao, Qiong
Li, Chang
Wang, Zhenfang
Xue, Chao
Fan, Yuwei
Xi, Wei
He, Xiaodong
Computation and Language
Artificial Intelligence
I.2.7
Large Reasoning Models (LRMs) demonstrate strong performance on complex tasks but often suffer from excessive verbosity, known as "overthinking." Existing solutions via reinforcement learning (RL) typically penalize generated tokens to promote conciseness. However, these methods encounter two challenges: responses with fewer tokens do not always correspond to fewer reasoning steps, and models may develop hacking behavior in later stages of training by discarding reasoning steps to minimize token usage. In this work, we introduce \textbf{Step Pruner (SP)}, an RL framework that steers LRMs toward more efficient reasoning by favoring compact reasoning steps. Our step-aware reward function prioritizes correctness while imposing penalties for redundant steps, and withholds rewards for incorrect responses to prevent the reinforcement of erroneous reasoning. Moreover, we propose a dynamic stopping mechanism: when the model's output no longer shortens, training is halted to prevent hacking behavior caused by the merging of steps. Extensive experiments across four reasoning benchmarks demonstrate that SP achieves state-of-the-art accuracy while significantly reducing response length. For instance, on AIME24, SP reduces token usage by \textbf{69.7\%}.
title Beyond Token Length: Step Pruner for Efficient and Accurate Reasoning in Large Language Models
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2510.03805