Your Group-Relative Advantage Is Biased

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Fengkai, Chen, Zherui, Wang, Xiaohan, Lu, Xiaodong, Chai, Jiajun, Yin, Guojun, Lin, Wei, Ma, Shuai, Zhuang, Fuzhen, Wang, Deqing, Yang, Yaodong, Li, Jianxin, Ban, Yikun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908780496683008
author Yang, Fengkai
Chen, Zherui
Wang, Xiaohan
Lu, Xiaodong
Chai, Jiajun
Yin, Guojun
Lin, Wei
Ma, Shuai
Zhuang, Fuzhen
Wang, Deqing
Yang, Yaodong
Li, Jianxin
Ban, Yikun
author_facet Yang, Fengkai
Chen, Zherui
Wang, Xiaohan
Lu, Xiaodong
Chai, Jiajun
Yin, Guojun
Lin, Wei
Ma, Shuai
Zhuang, Fuzhen
Wang, Deqing
Yang, Yaodong
Li, Jianxin
Ban, Yikun
contents Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such as GRPO and its variants gaining broad adoption. These methods rely on group-relative advantage estimation to avoid learned critics, yet its theoretical properties remain poorly understood. In this work, we uncover a fundamental issue of group-based RL: the group-relative advantage estimator is inherently biased relative to the true (expected) advantage. We provide the first theoretical analysis showing that it systematically underestimates advantages for hard prompts and overestimates them for easy prompts, leading to imbalanced exploration and exploitation. To address this issue, we propose History-Aware Adaptive Difficulty Weighting (HA-DW), an adaptive reweighting scheme that adjusts advantage estimates based on an evolving difficulty anchor and training dynamics. Both theoretical analysis and experiments on five mathematical reasoning benchmarks demonstrate that HA-DW consistently improves performance when integrated into GRPO and its variants. Our results suggest that correcting biased advantage estimation is critical for robust and efficient RLVR training.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08521
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Your Group-Relative Advantage Is Biased
Yang, Fengkai
Chen, Zherui
Wang, Xiaohan
Lu, Xiaodong
Chai, Jiajun
Yin, Guojun
Lin, Wei
Ma, Shuai
Zhuang, Fuzhen
Wang, Deqing
Yang, Yaodong
Li, Jianxin
Ban, Yikun
Machine Learning
Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such as GRPO and its variants gaining broad adoption. These methods rely on group-relative advantage estimation to avoid learned critics, yet its theoretical properties remain poorly understood. In this work, we uncover a fundamental issue of group-based RL: the group-relative advantage estimator is inherently biased relative to the true (expected) advantage. We provide the first theoretical analysis showing that it systematically underestimates advantages for hard prompts and overestimates them for easy prompts, leading to imbalanced exploration and exploitation. To address this issue, we propose History-Aware Adaptive Difficulty Weighting (HA-DW), an adaptive reweighting scheme that adjusts advantage estimates based on an evolving difficulty anchor and training dynamics. Both theoretical analysis and experiments on five mathematical reasoning benchmarks demonstrate that HA-DW consistently improves performance when integrated into GRPO and its variants. Our results suggest that correcting biased advantage estimation is critical for robust and efficient RLVR training.
title Your Group-Relative Advantage Is Biased
topic Machine Learning
url https://arxiv.org/abs/2601.08521