RewardAnything: Generalizable Principle-Following Reward Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yu, Zhuohao, Zeng, Jiali, Gu, Weizheng, Wang, Yidong, Wang, Jindong, Meng, Fandong, Zhou, Jie, Zhang, Yue, Zhang, Shikun, Ye, Wei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915373070155776
author Yu, Zhuohao
Zeng, Jiali
Gu, Weizheng
Wang, Yidong
Wang, Jindong
Meng, Fandong
Zhou, Jie
Zhang, Yue
Zhang, Shikun
Ye, Wei
author_facet Yu, Zhuohao
Zeng, Jiali
Gu, Weizheng
Wang, Yidong
Wang, Jindong
Meng, Fandong
Zhou, Jie
Zhang, Yue
Zhang, Shikun
Ye, Wei
contents Reward Models, essential for guiding Large Language Model optimization, are typically trained on fixed preference datasets, resulting in rigid alignment to single, implicit preference distributions. This prevents adaptation to diverse real-world needs-from conciseness in one task to detailed explanations in another. The standard practice of collecting task-specific preference data and retraining reward models is resource-intensive, often producing biased rewards, and limits practical application. We introduce generalizable, principle-following reward models. We propose that RMs should understand and adhere to dynamically provided natural language specifications of reward principles, similar to instruction-following in LLMs. To measure this capability, we develop RABench, a comprehensive benchmark for RMs focusing on generalization across diverse principles. Evaluations on RABench reveal poor generalization of current RMs. As a solution, we present RewardAnything, a novel RM designed and trained to explicitly follow natural language principles. We achieve SotA performance with RewardAnything in traditional RM benchmark simply by specifying a well-defined principle, and results on RABench show we excel in adapting to novel principles without retraining. Furthermore, RewardAnything integrates seamlessly with existing RLHF methods and we show by a case study on how to automatically and efficiently align LLMs with only natural language principles.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03637
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RewardAnything: Generalizable Principle-Following Reward Models
Yu, Zhuohao
Zeng, Jiali
Gu, Weizheng
Wang, Yidong
Wang, Jindong
Meng, Fandong
Zhou, Jie
Zhang, Yue
Zhang, Shikun
Ye, Wei
Computation and Language
Artificial Intelligence
Machine Learning
Reward Models, essential for guiding Large Language Model optimization, are typically trained on fixed preference datasets, resulting in rigid alignment to single, implicit preference distributions. This prevents adaptation to diverse real-world needs-from conciseness in one task to detailed explanations in another. The standard practice of collecting task-specific preference data and retraining reward models is resource-intensive, often producing biased rewards, and limits practical application. We introduce generalizable, principle-following reward models. We propose that RMs should understand and adhere to dynamically provided natural language specifications of reward principles, similar to instruction-following in LLMs. To measure this capability, we develop RABench, a comprehensive benchmark for RMs focusing on generalization across diverse principles. Evaluations on RABench reveal poor generalization of current RMs. As a solution, we present RewardAnything, a novel RM designed and trained to explicitly follow natural language principles. We achieve SotA performance with RewardAnything in traditional RM benchmark simply by specifying a well-defined principle, and results on RABench show we excel in adapting to novel principles without retraining. Furthermore, RewardAnything integrates seamlessly with existing RLHF methods and we show by a case study on how to automatically and efficiently align LLMs with only natural language principles.
title RewardAnything: Generalizable Principle-Following Reward Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.03637