Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Xueru, Lou, Jie, Li, Zichao, Lu, Yaojie, Yu, Xing, Ji, Yuqiu, Xu, Guohai, Lin, Hongyu, He, Ben, Han, Xianpei, Sun, Le, Zhang, Debing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916758449815552
author Wen, Xueru
Lou, Jie
Li, Zichao
Lu, Yaojie
Yu, Xing
Ji, Yuqiu
Xu, Guohai
Lin, Hongyu
He, Ben
Han, Xianpei
Sun, Le
Zhang, Debing
author_facet Wen, Xueru
Lou, Jie
Li, Zichao
Lu, Yaojie
Yu, Xing
Ji, Yuqiu
Xu, Guohai
Lin, Hongyu
He, Ben
Han, Xianpei
Sun, Le
Zhang, Debing
contents Reward models (RMs) are crucial for aligning large language models (LLMs) with human preferences. However, most RM research is centered on English and relies heavily on synthetic resources, which leads to limited and less reliable datasets and benchmarks for Chinese. To address this gap, we introduce CheemsBench, a fully human-annotated RM evaluation benchmark within Chinese contexts, and CheemsPreference, a large-scale and diverse preference dataset annotated through human-machine collaboration to support Chinese RM training. We systematically evaluate open-source discriminative and generative RMs on CheemsBench and observe significant limitations in their ability to capture human preferences in Chinese scenarios. Additionally, based on CheemsPreference, we construct an RM that achieves state-of-the-art performance on CheemsBench, demonstrating the necessity of human supervision in RM training. Our findings reveal that scaled AI-generated data struggles to fully capture human preferences, emphasizing the importance of high-quality human supervision in RM development.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17173
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch
Wen, Xueru
Lou, Jie
Li, Zichao
Lu, Yaojie
Yu, Xing
Ji, Yuqiu
Xu, Guohai
Lin, Hongyu
He, Ben
Han, Xianpei
Sun, Le
Zhang, Debing
Computation and Language
Artificial Intelligence
Reward models (RMs) are crucial for aligning large language models (LLMs) with human preferences. However, most RM research is centered on English and relies heavily on synthetic resources, which leads to limited and less reliable datasets and benchmarks for Chinese. To address this gap, we introduce CheemsBench, a fully human-annotated RM evaluation benchmark within Chinese contexts, and CheemsPreference, a large-scale and diverse preference dataset annotated through human-machine collaboration to support Chinese RM training. We systematically evaluate open-source discriminative and generative RMs on CheemsBench and observe significant limitations in their ability to capture human preferences in Chinese scenarios. Additionally, based on CheemsPreference, we construct an RM that achieves state-of-the-art performance on CheemsBench, demonstrating the necessity of human supervision in RM training. Our findings reveal that scaled AI-generated data struggles to fully capture human preferences, emphasizing the importance of high-quality human supervision in RM development.
title Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.17173