CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Qingbin, Xue, Rongkun, Wang, Jie, Zhou, Ming, Li, Zhi, Ji, Xiaofeng, Wang, Yongqi, Liu, Miao, Yang, Zheming, Qiu, Minghui, Yang, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909750592012288
author Li, Qingbin
Xue, Rongkun
Wang, Jie
Zhou, Ming
Li, Zhi
Ji, Xiaofeng
Wang, Yongqi
Liu, Miao
Yang, Zheming
Qiu, Minghui
Yang, Jing
author_facet Li, Qingbin
Xue, Rongkun
Wang, Jie
Zhou, Ming
Li, Zhi
Ji, Xiaofeng
Wang, Yongqi
Liu, Miao
Yang, Zheming
Qiu, Minghui
Yang, Jing
contents Recent advances in Reinforcement Learning with Verified Reward (RLVR) have driven the emergence of more sophisticated cognitive behaviors in large language models (LLMs), thereby enhancing their reasoning capabilities. However, in prior RLVR pipelines, the repeated use of static initial-state sampling drawn exactly from the dataset distribution during each sampling phase produced overly deterministic, low diversity model behavior, which manifested as rapid entropy collapse and hindered sustained performance gains during prolonged training. To address this issue, we introduce CURE (Critical-token-gUided Re concatenation for Entropy-collapse prevention), a two-stage framework that balances exploration and exploitation. Specifically, in the first stage, to deliberately steer the model toward novel yet coherent contexts, we re-generate at high-entropy critical tokens and jointly optimize the original and the branched trajectories. The further comparison with vanilla DAPO shows that the regeneration process achieves a better performance on math reasoning tasks while sustaining a high-level entropy degree for exploration. In the second stage, we continue training with static initial-state sampling by DAPO, intentionally placing the model in a familiar state to gradually strengthen exploitation. Extensive experiments on Qwen-2.5-Math-7B show that, compared to other RLVR methods, CURE achieves a 5% performance gain across six math benchmarks, establishing state-of-the-art performance in both entropy and accuracy. A series of experiments further validate the effectiveness of our approach. Code is available at https://github.com/bytedance/CURE.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11016
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
Li, Qingbin
Xue, Rongkun
Wang, Jie
Zhou, Ming
Li, Zhi
Ji, Xiaofeng
Wang, Yongqi
Liu, Miao
Yang, Zheming
Qiu, Minghui
Yang, Jing
Machine Learning
Artificial Intelligence
Recent advances in Reinforcement Learning with Verified Reward (RLVR) have driven the emergence of more sophisticated cognitive behaviors in large language models (LLMs), thereby enhancing their reasoning capabilities. However, in prior RLVR pipelines, the repeated use of static initial-state sampling drawn exactly from the dataset distribution during each sampling phase produced overly deterministic, low diversity model behavior, which manifested as rapid entropy collapse and hindered sustained performance gains during prolonged training. To address this issue, we introduce CURE (Critical-token-gUided Re concatenation for Entropy-collapse prevention), a two-stage framework that balances exploration and exploitation. Specifically, in the first stage, to deliberately steer the model toward novel yet coherent contexts, we re-generate at high-entropy critical tokens and jointly optimize the original and the branched trajectories. The further comparison with vanilla DAPO shows that the regeneration process achieves a better performance on math reasoning tasks while sustaining a high-level entropy degree for exploration. In the second stage, we continue training with static initial-state sampling by DAPO, intentionally placing the model in a familiar state to gradually strengthen exploitation. Extensive experiments on Qwen-2.5-Math-7B show that, compared to other RLVR methods, CURE achieves a 5% performance gain across six math benchmarks, establishing state-of-the-art performance in both entropy and accuracy. A series of experiments further validate the effectiveness of our approach. Code is available at https://github.com/bytedance/CURE.
title CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.11016