R^3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Zhizheng, Zhao, Kang, Xu, Weikai, Lin, Xinkui, Liu, Wei, Luan, Jian, Shang, Shuo, Han, Peng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915758536130560
author Jiang, Zhizheng
Zhao, Kang
Xu, Weikai
Lin, Xinkui
Liu, Wei
Luan, Jian
Shang, Shuo
Han, Peng
author_facet Jiang, Zhizheng
Zhao, Kang
Xu, Weikai
Lin, Xinkui
Liu, Wei
Luan, Jian
Shang, Shuo
Han, Peng
contents Large reasoning models (LRMs) aim to solve diverse and complex problems through structured reasoning. Recent advances in group-based policy optimization methods have shown promise in enabling stable advantage estimation without reliance on process-level annotations. However, these methods rely on advantage gaps induced by high-quality samples within the same batch, which makes the training process fragile and inefficient when intra-group advantages collapse under challenging tasks. To address these problems, we propose a reinforcement learning mechanism named \emph{\textbf{R^3}} that along three directions: (1) a \emph{cross-context \underline{\textbf{R}}eplay} strategy that maintains the intra-group advantage by recalling valuable examples from historical trajectories of the same query, (2) an \emph{in-context self-\underline{\textbf{R}}eflection} mechanism enabling models to refine outputs by leveraging past failures, and (3) a \emph{structural entropy \underline{\textbf{R}}anking reward}, which assigns relative rewards to truncated or failed samples by ranking responses based on token-level entropy patterns, capturing both local exploration and global stability. We implement our method on Deepseek-R1-Distill-Qwen-1.5B and train it on the DeepscaleR-40k in the math domain. Experiments demonstrate our method achieves SoTA performance on several math benchmarks, representing significant improvements and fewer reasoning tokens over the base models. Code and model will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19620
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle R^3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning
Jiang, Zhizheng
Zhao, Kang
Xu, Weikai
Lin, Xinkui
Liu, Wei
Luan, Jian
Shang, Shuo
Han, Peng
Machine Learning
Artificial Intelligence
Large reasoning models (LRMs) aim to solve diverse and complex problems through structured reasoning. Recent advances in group-based policy optimization methods have shown promise in enabling stable advantage estimation without reliance on process-level annotations. However, these methods rely on advantage gaps induced by high-quality samples within the same batch, which makes the training process fragile and inefficient when intra-group advantages collapse under challenging tasks. To address these problems, we propose a reinforcement learning mechanism named \emph{\textbf{R^3}} that along three directions: (1) a \emph{cross-context \underline{\textbf{R}}eplay} strategy that maintains the intra-group advantage by recalling valuable examples from historical trajectories of the same query, (2) an \emph{in-context self-\underline{\textbf{R}}eflection} mechanism enabling models to refine outputs by leveraging past failures, and (3) a \emph{structural entropy \underline{\textbf{R}}anking reward}, which assigns relative rewards to truncated or failed samples by ranking responses based on token-level entropy patterns, capturing both local exploration and global stability. We implement our method on Deepseek-R1-Distill-Qwen-1.5B and train it on the DeepscaleR-40k in the math domain. Experiments demonstrate our method achieves SoTA performance on several math benchmarks, representing significant improvements and fewer reasoning tokens over the base models. Code and model will be released.
title R^3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2601.19620