Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jia, Ruipeng, Yang, Yunyi, Gai, Yongbo, Luo, Kai, Huang, Shihao, Lin, Jianhe, Jiang, Xiaoxi, Jiang, Guanjun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910999806738432
author Jia, Ruipeng
Yang, Yunyi
Gai, Yongbo
Luo, Kai
Huang, Shihao
Lin, Jianhe
Jiang, Xiaoxi
Jiang, Guanjun
author_facet Jia, Ruipeng
Yang, Yunyi
Gai, Yongbo
Luo, Kai
Huang, Shihao
Lin, Jianhe
Jiang, Xiaoxi
Jiang, Guanjun
contents Reinforcement learning with verifiable rewards (RLVR) has enabled large language models (LLMs) to achieve remarkable breakthroughs in reasoning tasks with objective ground-truth answers, such as mathematics and code generation. However, a significant gap remains for non-verifiable tasks, like creative writing and open-ended dialogue, where quality assessment is inherently subjective and lacks definitive references. Existing approaches for these domains often rely on scalar reward models trained with human preferences, which suffer from limited generalization and are prone to reward hacking, such as over-explanation and length bias. In this work, we propose a unified RLVR-based training paradigm that bridges the gap between non-verifiable tasks and verifiable rewards. We introduce a writing-principle-based pairwise Generative Reward Model (GenRM) and a novel Bootstrapped Relative Policy Optimization (BRPO) algorithm. The pairwise writing GenRM leverages self-principled critique to transform subjective assessments into reliable, verifiable rewards, while BRPO enables dynamic, reference-free pairwise comparison by leveraging a bootstrapped response as temporary reference from within group rollouts during RL training. Our approach empowers LLMs to develop robust writing capabilities without supervised fine-tuning, as demonstrated by Writing-Zero, which shows consistent improvement and strong resistance to reward hacking compared to scalar reward baselines. Furthermore, our method achieves competitive results on both in-house and open-source writing benchmarks. Our findings suggest the potential to unify rule-based, reference-based, and reference-free reward modeling under the RLVR framework, thus paving the way for a comprehensive and scalable RL training paradigm applicable across all language tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00103
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
Jia, Ruipeng
Yang, Yunyi
Gai, Yongbo
Luo, Kai
Huang, Shihao
Lin, Jianhe
Jiang, Xiaoxi
Jiang, Guanjun
Computation and Language
Reinforcement learning with verifiable rewards (RLVR) has enabled large language models (LLMs) to achieve remarkable breakthroughs in reasoning tasks with objective ground-truth answers, such as mathematics and code generation. However, a significant gap remains for non-verifiable tasks, like creative writing and open-ended dialogue, where quality assessment is inherently subjective and lacks definitive references. Existing approaches for these domains often rely on scalar reward models trained with human preferences, which suffer from limited generalization and are prone to reward hacking, such as over-explanation and length bias. In this work, we propose a unified RLVR-based training paradigm that bridges the gap between non-verifiable tasks and verifiable rewards. We introduce a writing-principle-based pairwise Generative Reward Model (GenRM) and a novel Bootstrapped Relative Policy Optimization (BRPO) algorithm. The pairwise writing GenRM leverages self-principled critique to transform subjective assessments into reliable, verifiable rewards, while BRPO enables dynamic, reference-free pairwise comparison by leveraging a bootstrapped response as temporary reference from within group rollouts during RL training. Our approach empowers LLMs to develop robust writing capabilities without supervised fine-tuning, as demonstrated by Writing-Zero, which shows consistent improvement and strong resistance to reward hacking compared to scalar reward baselines. Furthermore, our method achieves competitive results on both in-house and open-source writing benchmarks. Our findings suggest the potential to unify rule-based, reference-based, and reference-free reward modeling under the RLVR framework, thus paving the way for a comprehensive and scalable RL training paradigm applicable across all language tasks.
title Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
topic Computation and Language
url https://arxiv.org/abs/2506.00103