Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Yi, Yu, Dian, Song, Linfeng, Li, Juntao, Mi, Haitao, Tu, Zhaopeng, Zhang, Min, Yu, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916669570416640
author Su, Yi
Yu, Dian
Song, Linfeng
Li, Juntao
Mi, Haitao
Tu, Zhaopeng
Zhang, Min
Yu, Dong
author_facet Su, Yi
Yu, Dian
Song, Linfeng
Li, Juntao
Mi, Haitao
Tu, Zhaopeng
Zhang, Min
Yu, Dong
contents Reinforcement learning with verifiable rewards (RLVR) has demonstrated significant success in enhancing mathematical reasoning and coding performance of large language models (LLMs), especially when structured reference answers are accessible for verification. However, its extension to broader, less structured domains remains unexplored. In this work, we investigate the effectiveness and scalability of RLVR across diverse real-world domains including medicine, chemistry, psychology, economics, and education, where structured reference answers are typically unavailable. We reveal that binary verification judgments on broad-domain tasks exhibit high consistency across various LLMs provided expert-written reference answers exist. Motivated by this finding, we utilize a generative scoring technique that yields soft, model-based reward signals to overcome limitations posed by binary verifications, especially in free-form, unstructured answer scenarios. We further demonstrate the feasibility of training cross-domain generative reward models using relatively small (7B) LLMs without the need for extensive domain-specific annotation. Through comprehensive experiments, our RLVR framework establishes clear performance gains, significantly outperforming state-of-the-art open-source aligned models such as Qwen2.5-72B and DeepSeek-R1-Distill-Qwen-32B across domains in free-form settings. Our approach notably enhances the robustness, flexibility, and scalability of RLVR, representing a substantial step towards practical reinforcement learning applications in complex, noisy-label scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23829
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Su, Yi
Yu, Dian
Song, Linfeng
Li, Juntao
Mi, Haitao
Tu, Zhaopeng
Zhang, Min
Yu, Dong
Computation and Language
Reinforcement learning with verifiable rewards (RLVR) has demonstrated significant success in enhancing mathematical reasoning and coding performance of large language models (LLMs), especially when structured reference answers are accessible for verification. However, its extension to broader, less structured domains remains unexplored. In this work, we investigate the effectiveness and scalability of RLVR across diverse real-world domains including medicine, chemistry, psychology, economics, and education, where structured reference answers are typically unavailable. We reveal that binary verification judgments on broad-domain tasks exhibit high consistency across various LLMs provided expert-written reference answers exist. Motivated by this finding, we utilize a generative scoring technique that yields soft, model-based reward signals to overcome limitations posed by binary verifications, especially in free-form, unstructured answer scenarios. We further demonstrate the feasibility of training cross-domain generative reward models using relatively small (7B) LLMs without the need for extensive domain-specific annotation. Through comprehensive experiments, our RLVR framework establishes clear performance gains, significantly outperforming state-of-the-art open-source aligned models such as Qwen2.5-72B and DeepSeek-R1-Distill-Qwen-32B across domains in free-form settings. Our approach notably enhances the robustness, flexibility, and scalability of RLVR, representing a substantial step towards practical reinforcement learning applications in complex, noisy-label scenarios.
title Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
topic Computation and Language
url https://arxiv.org/abs/2503.23829