Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mansouri, Omar El, Izzati, Fathinah Asma, Seddik, Mohamed El Amine, Lahlou, Salem
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913145561284608
author Mansouri, Omar El
Izzati, Fathinah Asma
Seddik, Mohamed El Amine
Lahlou, Salem
author_facet Mansouri, Omar El
Izzati, Fathinah Asma
Seddik, Mohamed El Amine
Lahlou, Salem
contents Reinforcement learning from human feedback (RLHF) or verifiable rewards (RLVR), the standard paradigm for aligning LLMs or building recent SOTA reasoning models, is highly sensitive to noise from inconsistent or erroneous rewards. Yet, the interaction between such noise and widely used group-based policy optimization methods remains underexplored. We introduce a noise-robust Group Relative Policy Optimization (GRPO) and Done Right GRPO (Dr.GRPO) framework that explicitly models reward corruption as Bernoulli noise. Our method applies noise correction after estimating reward flip probabilities to debias the learning signal, yielding provably unbiased gradient estimates. Theoretical analysis shows that group-based methods inherently mitigate individual-level noise, and our correction strategy amplifies this robustness. Empirically, we observe consistent improvements across math and code tasks when applying our noise correction to standard reward model usage, with particular gains of up to 6.7 percentage points in accuracy on math tasks and 1.5 on code tasks under realistic reward model conditions. This work bridges label-noise correction from supervised learning with modern RLHF, offering both theoretical insights and a practical algorithm for noisy real-world deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18924
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
Mansouri, Omar El
Izzati, Fathinah Asma
Seddik, Mohamed El Amine
Lahlou, Salem
Machine Learning
Artificial Intelligence
Reinforcement learning from human feedback (RLHF) or verifiable rewards (RLVR), the standard paradigm for aligning LLMs or building recent SOTA reasoning models, is highly sensitive to noise from inconsistent or erroneous rewards. Yet, the interaction between such noise and widely used group-based policy optimization methods remains underexplored. We introduce a noise-robust Group Relative Policy Optimization (GRPO) and Done Right GRPO (Dr.GRPO) framework that explicitly models reward corruption as Bernoulli noise. Our method applies noise correction after estimating reward flip probabilities to debias the learning signal, yielding provably unbiased gradient estimates. Theoretical analysis shows that group-based methods inherently mitigate individual-level noise, and our correction strategy amplifies this robustness. Empirically, we observe consistent improvements across math and code tasks when applying our noise correction to standard reward model usage, with particular gains of up to 6.7 percentage points in accuracy on math tasks and 1.5 on code tasks under realistic reward model conditions. This work bridges label-noise correction from supervised learning with modern RLHF, offering both theoretical insights and a practical algorithm for noisy real-world deployment.
title Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.18924