Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Egashira, Kazuki, Vero, Mark, Dekoninck, Jasper, Dorner, Florian E., Staab, Robin, Vechev, Martin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918482209144832
author Egashira, Kazuki
Vero, Mark
Dekoninck, Jasper
Dorner, Florian E.
Staab, Robin
Vechev, Martin
author_facet Egashira, Kazuki
Vero, Mark
Dekoninck, Jasper
Dorner, Florian E.
Staab, Robin
Vechev, Martin
contents Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs). While RLVR is designed for tasks with verifiable ground-truth answers, real-world verifiers (e.g., static code checkers) can introduce errors into the reward signal. Prior analyses have largely treated such errors as random and independent across samples, concluding that errors merely slow training with limited effect on final performance. However, practical verifiers tend to exhibit systematic errors. This introduces a risk of models learning unwanted consistent behavior from a structurally incorrect reward signal. In this work, we study the impact of such systematic verification errors on RLVR. Through controlled experiments on arithmetic tasks, we show that systematic false negatives lead to similar effects as random noise. On the other hand, systematic false positives can cause a wide range of behaviors from sub-optimal plateaus to performance collapse. Crucially, these outcomes are not determined by the overall error rate but by the specific pattern of introduced errors, making pre-hoc mitigation difficult. Our results show that, in contrast to prior conclusions, realistic verification errors can critically shape RLVR outcomes and that verifier quality has to be understood beyond its sample-level error rate.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02909
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR
Egashira, Kazuki
Vero, Mark
Dekoninck, Jasper
Dorner, Florian E.
Staab, Robin
Vechev, Martin
Machine Learning
Artificial Intelligence
Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs). While RLVR is designed for tasks with verifiable ground-truth answers, real-world verifiers (e.g., static code checkers) can introduce errors into the reward signal. Prior analyses have largely treated such errors as random and independent across samples, concluding that errors merely slow training with limited effect on final performance. However, practical verifiers tend to exhibit systematic errors. This introduces a risk of models learning unwanted consistent behavior from a structurally incorrect reward signal. In this work, we study the impact of such systematic verification errors on RLVR. Through controlled experiments on arithmetic tasks, we show that systematic false negatives lead to similar effects as random noise. On the other hand, systematic false positives can cause a wide range of behaviors from sub-optimal plateaus to performance collapse. Crucially, these outcomes are not determined by the overall error rate but by the specific pattern of introduced errors, making pre-hoc mitigation difficult. Our results show that, in contrast to prior conclusions, realistic verification errors can critically shape RLVR outcomes and that verifier quality has to be understood beyond its sample-level error rate.
title Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.02909