Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zongxia, Chang, Yapei, Zhou, Yuhang, Wu, Xiyang, Liang, Zichao, Sung, Yoo Yeon, Boyd-Graber, Jordan Lee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918062229291008
author Li, Zongxia
Chang, Yapei
Zhou, Yuhang
Wu, Xiyang
Liang, Zichao
Sung, Yoo Yeon
Boyd-Graber, Jordan Lee
author_facet Li, Zongxia
Chang, Yapei
Zhou, Yuhang
Wu, Xiyang
Liang, Zichao
Sung, Yoo Yeon
Boyd-Graber, Jordan Lee
contents Evaluating open-ended long-form generation is challenging because it is hard to define what clearly separates good from bad outputs. Existing methods often miss key aspects like coherence, style, or relevance, or are biased by pretraining data, making open-ended long-form evaluation an underexplored problem. To address this gap, we propose PrefBERT, a scoring model for evaluating open-ended long-form generation in GRPO and guiding its training with distinct rewards for good and bad outputs. Trained on two response evaluation datasets with diverse long-form styles and Likert-rated quality, PrefBERT effectively supports GRPO by offering better semantic reward feedback than traditional metrics ROUGE-L and BERTScore do. Through comprehensive evaluations, including LLM-as-a-judge, human ratings, and qualitative analysis, we show that PrefBERT, trained on multi-sentence and paragraph-length responses, remains reliable across varied long passages and aligns well with the verifiable rewards GRPO needs. Human evaluations confirm that using PrefBERT as the reward signal to train policy models yields responses better aligned with human preferences than those trained with traditional metrics. Our code is available at https://github.com/zli12321/long_form_rl.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15068
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
Li, Zongxia
Chang, Yapei
Zhou, Yuhang
Wu, Xiyang
Liang, Zichao
Sung, Yoo Yeon
Boyd-Graber, Jordan Lee
Computation and Language
Machine Learning
Evaluating open-ended long-form generation is challenging because it is hard to define what clearly separates good from bad outputs. Existing methods often miss key aspects like coherence, style, or relevance, or are biased by pretraining data, making open-ended long-form evaluation an underexplored problem. To address this gap, we propose PrefBERT, a scoring model for evaluating open-ended long-form generation in GRPO and guiding its training with distinct rewards for good and bad outputs. Trained on two response evaluation datasets with diverse long-form styles and Likert-rated quality, PrefBERT effectively supports GRPO by offering better semantic reward feedback than traditional metrics ROUGE-L and BERTScore do. Through comprehensive evaluations, including LLM-as-a-judge, human ratings, and qualitative analysis, we show that PrefBERT, trained on multi-sentence and paragraph-length responses, remains reliable across varied long passages and aligns well with the verifiable rewards GRPO needs. Human evaluations confirm that using PrefBERT as the reward signal to train policy models yields responses better aligned with human preferences than those trained with traditional metrics. Our code is available at https://github.com/zli12321/long_form_rl.
title Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.15068