RLPR: Extrapolating RLVR to General Domains without Verifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Tianyu, Ji, Bo, Wang, Shouli, Yao, Shu, Wang, Zefan, Cui, Ganqu, Yuan, Lifan, Ding, Ning, Yao, Yuan, Liu, Zhiyuan, Sun, Maosong, Chua, Tat-Seng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908416710017024
author Yu, Tianyu
Ji, Bo
Wang, Shouli
Yao, Shu
Wang, Zefan
Cui, Ganqu
Yuan, Lifan
Ding, Ning
Yao, Yuan
Liu, Zhiyuan
Sun, Maosong
Chua, Tat-Seng
author_facet Yu, Tianyu
Ji, Bo
Wang, Shouli
Yao, Shu
Wang, Zefan
Cui, Ganqu
Yuan, Lifan
Ding, Ning
Yao, Yuan
Liu, Zhiyuan
Sun, Maosong
Chua, Tat-Seng
contents Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates promising potential in advancing the reasoning capabilities of LLMs. However, its success remains largely confined to mathematical and code domains. This primary limitation stems from the heavy reliance on domain-specific verifiers, which results in prohibitive complexity and limited scalability. To address the challenge, our key observation is that LLM's intrinsic probability of generating a correct free-form answer directly indicates its own evaluation of the reasoning reward (i.e., how well the reasoning process leads to the correct answer). Building on this insight, we propose RLPR, a simple verifier-free framework that extrapolates RLVR to broader general domains. RLPR uses the LLM's own token probability scores for reference answers as the reward signal and maximizes the expected reward during training. We find that addressing the high variance of this noisy probability reward is crucial to make it work, and propose prob-to-reward and stabilizing methods to ensure a precise and stable reward from LLM intrinsic probabilities. Comprehensive experiments in four general-domain benchmarks and three mathematical benchmarks show that RLPR consistently improves reasoning capabilities in both areas for Gemma, Llama, and Qwen based models. Notably, RLPR outperforms concurrent VeriFree by 7.6 points on TheoremQA and 7.5 points on Minerva, and even surpasses strong verifier-model-dependent approaches General-Reasoner by 1.6 average points across seven benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18254
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RLPR: Extrapolating RLVR to General Domains without Verifiers
Yu, Tianyu
Ji, Bo
Wang, Shouli
Yao, Shu
Wang, Zefan
Cui, Ganqu
Yuan, Lifan
Ding, Ning
Yao, Yuan
Liu, Zhiyuan
Sun, Maosong
Chua, Tat-Seng
Machine Learning
Artificial Intelligence
Computation and Language
Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates promising potential in advancing the reasoning capabilities of LLMs. However, its success remains largely confined to mathematical and code domains. This primary limitation stems from the heavy reliance on domain-specific verifiers, which results in prohibitive complexity and limited scalability. To address the challenge, our key observation is that LLM's intrinsic probability of generating a correct free-form answer directly indicates its own evaluation of the reasoning reward (i.e., how well the reasoning process leads to the correct answer). Building on this insight, we propose RLPR, a simple verifier-free framework that extrapolates RLVR to broader general domains. RLPR uses the LLM's own token probability scores for reference answers as the reward signal and maximizes the expected reward during training. We find that addressing the high variance of this noisy probability reward is crucial to make it work, and propose prob-to-reward and stabilizing methods to ensure a precise and stable reward from LLM intrinsic probabilities. Comprehensive experiments in four general-domain benchmarks and three mathematical benchmarks show that RLPR consistently improves reasoning capabilities in both areas for Gemma, Llama, and Qwen based models. Notably, RLPR outperforms concurrent VeriFree by 7.6 points on TheoremQA and 7.5 points on Minerva, and even surpasses strong verifier-model-dependent approaches General-Reasoner by 1.6 average points across seven benchmarks.
title RLPR: Extrapolating RLVR to General Domains without Verifiers
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.18254