WildReward: Learning Reward Models from In-the-Wild Human Interactions

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Peng, Hao, Qi, Yunjia, Wang, Xiaozhi, Yao, Zijun, Hou, Lei, Li, Juanzi
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914315825577984
author Peng, Hao
Qi, Yunjia
Wang, Xiaozhi
Yao, Zijun
Hou, Lei
Li, Juanzi
author_facet Peng, Hao
Qi, Yunjia
Wang, Xiaozhi
Yao, Zijun
Hou, Lei
Li, Juanzi
contents Reward models (RMs) are crucial for the training of large language models (LLMs), yet they typically rely on large-scale human-annotated preference pairs. With the widespread deployment of LLMs, in-the-wild interactions have emerged as a rich source of implicit reward signals. This raises the question: Can we develop reward models directly from in-the-wild interactions? In this work, we explore this possibility by adopting WildChat as an interaction source and proposing a pipeline to extract reliable human feedback, yielding 186k high-quality instances for training WildReward via ordinal regression directly on user feedback without preference pairs. Extensive experiments demonstrate that WildReward achieves comparable or even superior performance compared to conventional reward models, with improved calibration and cross-sample consistency. We also observe that WildReward benefits directly from user diversity, where more users yield stronger reward models. Finally, we apply WildReward to online DPO training and observe significant improvements across various tasks. Code and data are released at https://github.com/THU-KEG/WildReward.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08829
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WildReward: Learning Reward Models from In-the-Wild Human Interactions
Peng, Hao
Qi, Yunjia
Wang, Xiaozhi
Yao, Zijun
Hou, Lei
Li, Juanzi
Computation and Language
Artificial Intelligence
Reward models (RMs) are crucial for the training of large language models (LLMs), yet they typically rely on large-scale human-annotated preference pairs. With the widespread deployment of LLMs, in-the-wild interactions have emerged as a rich source of implicit reward signals. This raises the question: Can we develop reward models directly from in-the-wild interactions? In this work, we explore this possibility by adopting WildChat as an interaction source and proposing a pipeline to extract reliable human feedback, yielding 186k high-quality instances for training WildReward via ordinal regression directly on user feedback without preference pairs. Extensive experiments demonstrate that WildReward achieves comparable or even superior performance compared to conventional reward models, with improved calibration and cross-sample consistency. We also observe that WildReward benefits directly from user diversity, where more users yield stronger reward models. Finally, we apply WildReward to online DPO training and observe significant improvements across various tasks. Code and data are released at https://github.com/THU-KEG/WildReward.
title WildReward: Learning Reward Models from In-the-Wild Human Interactions
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.08829