Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Benteng, Wang, Weida, Zhang, Shufei, Lin, Mingbao, Zhang, Min
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917417909747712
author Chen, Benteng
Wang, Weida
Zhang, Shufei
Lin, Mingbao
Zhang, Min
author_facet Chen, Benteng
Wang, Weida
Zhang, Shufei
Lin, Mingbao
Zhang, Min
contents Large reasoning models that use long chain-of-thought excel at problem-solving yet waste compute on redundant checks. Curbing this overthinking is hard: training-time length penalties can cripple ability, while inference-time early-exit adds system overhead. To bridge this gap, we propose Step-GRPO, a novel post-training framework that internalizes dynamic early-exit capabilities directly into the model. Step-GRPO shifts the optimization objective from raw tokens to semantic steps by utilizing linguistic markers to structure reasoning. We introduce a Dynamic Truncated Rollout mechanism that exposes the model to concise high-confidence trajectories during exploration, synergized with a Step-Aware Relative Reward that dynamically penalizes redundancy based on group-level baselines. Extensive experiments across three model sizes on diverse benchmarks demonstrate that Step-GRPO achieves a superior accuracy-efficiency trade-off. On Qwen3-8B, our method reduces token consumption by 32.0\% compared to the vanilla model while avoiding the accuracy degradation observed in traditional length-penalty methods.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16890
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
Chen, Benteng
Wang, Weida
Zhang, Shufei
Lin, Mingbao
Zhang, Min
Artificial Intelligence
Large reasoning models that use long chain-of-thought excel at problem-solving yet waste compute on redundant checks. Curbing this overthinking is hard: training-time length penalties can cripple ability, while inference-time early-exit adds system overhead. To bridge this gap, we propose Step-GRPO, a novel post-training framework that internalizes dynamic early-exit capabilities directly into the model. Step-GRPO shifts the optimization objective from raw tokens to semantic steps by utilizing linguistic markers to structure reasoning. We introduce a Dynamic Truncated Rollout mechanism that exposes the model to concise high-confidence trajectories during exploration, synergized with a Step-Aware Relative Reward that dynamically penalizes redundancy based on group-level baselines. Extensive experiments across three model sizes on diverse benchmarks demonstrate that Step-GRPO achieves a superior accuracy-efficiency trade-off. On Qwen3-8B, our method reduces token consumption by 32.0\% compared to the vanilla model while avoiding the accuracy degradation observed in traditional length-penalty methods.
title Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
topic Artificial Intelligence
url https://arxiv.org/abs/2604.16890