RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhilin, Zeng, Jiaqi, Delalleau, Olivier, Evans, Ellie, Egert, Daniel, Shin, Hoo-Chang, Soares, Felipe, Dong, Yi, Kuchaiev, Oleksii
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917503824822272
author Wang, Zhilin
Zeng, Jiaqi
Delalleau, Olivier
Evans, Ellie
Egert, Daniel
Shin, Hoo-Chang
Soares, Felipe
Dong, Yi
Kuchaiev, Oleksii
author_facet Wang, Zhilin
Zeng, Jiaqi
Delalleau, Olivier
Evans, Ellie
Egert, Daniel
Shin, Hoo-Chang
Soares, Felipe
Dong, Yi
Kuchaiev, Oleksii
contents Reinforcement Learning with Human Feedback (RLHF) and Reinforcement Learning with Verifiable Rewards (RLVR) are the main RL paradigms used in LLM post-training, each offering distinct advantages. However, RLHF struggles with interpretability and reward hacking because it relies on human judgments that usually lack explicit criteria, whereas RLVR is limited in scope by its focus on correctness-based verifiers. We propose Reinforcement Learning with Binary Flexible Feedback (RLBFF), which combines the versatility of human-driven preferences with the precision of rule-based verification, enabling reward models to capture nuanced aspects of response quality beyond mere correctness. RLBFF extracts principles that can be answered in a binary fashion (e.g. accuracy of information: yes, or code readability: no) from natural language feedback. Such principles can then be used to ground Reward Model training as an entailment task (response satisfies or does not satisfy an arbitrary principle). We show that Reward Models trained in this manner can outperform Bradley-Terry models when matched for data and achieve top performance on RM-Bench (86.2%) and JudgeBench (81.4%, #1 on leaderboard as of September 24, 2025). Additionally, users can specify principles of interest at inference time to customize the focus of our reward models, in contrast to Bradley-Terry models. Finally, we present a fully open source recipe (including data) to align Qwen3-32B using RLBFF and our Reward Model, to match or exceed the performance of o3-mini and DeepSeek R1 on general alignment benchmarks of MT-Bench, WildBench, and Arena Hard v2 (at <5% of the inference cost). Models: https://huggingface.co/collections/nvidia/reward-models-10-2025
format Preprint
id arxiv_https___arxiv_org_abs_2509_21319
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
Wang, Zhilin
Zeng, Jiaqi
Delalleau, Olivier
Evans, Ellie
Egert, Daniel
Shin, Hoo-Chang
Soares, Felipe
Dong, Yi
Kuchaiev, Oleksii
Computation and Language
Artificial Intelligence
Machine Learning
Reinforcement Learning with Human Feedback (RLHF) and Reinforcement Learning with Verifiable Rewards (RLVR) are the main RL paradigms used in LLM post-training, each offering distinct advantages. However, RLHF struggles with interpretability and reward hacking because it relies on human judgments that usually lack explicit criteria, whereas RLVR is limited in scope by its focus on correctness-based verifiers. We propose Reinforcement Learning with Binary Flexible Feedback (RLBFF), which combines the versatility of human-driven preferences with the precision of rule-based verification, enabling reward models to capture nuanced aspects of response quality beyond mere correctness. RLBFF extracts principles that can be answered in a binary fashion (e.g. accuracy of information: yes, or code readability: no) from natural language feedback. Such principles can then be used to ground Reward Model training as an entailment task (response satisfies or does not satisfy an arbitrary principle). We show that Reward Models trained in this manner can outperform Bradley-Terry models when matched for data and achieve top performance on RM-Bench (86.2%) and JudgeBench (81.4%, #1 on leaderboard as of September 24, 2025). Additionally, users can specify principles of interest at inference time to customize the focus of our reward models, in contrast to Bradley-Terry models. Finally, we present a fully open source recipe (including data) to align Qwen3-32B using RLBFF and our Reward Model, to match or exceed the performance of o3-mini and DeepSeek R1 on general alignment benchmarks of MT-Bench, WildBench, and Arena Hard v2 (at <5% of the inference cost). Models: https://huggingface.co/collections/nvidia/reward-models-10-2025
title RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.21319