HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Zhilin, Zeng, Jiaqi, Delalleau, Olivier, Shin, Hoo-Chang, Soares, Felipe, Bukharin, Alexander, Evans, Ellie, Dong, Yi, Kuchaiev, Oleksii
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911228751773696
author Wang, Zhilin
Zeng, Jiaqi
Delalleau, Olivier
Shin, Hoo-Chang
Soares, Felipe
Bukharin, Alexander
Evans, Ellie
Dong, Yi
Kuchaiev, Oleksii
author_facet Wang, Zhilin
Zeng, Jiaqi
Delalleau, Olivier
Shin, Hoo-Chang
Soares, Felipe
Bukharin, Alexander
Evans, Ellie
Dong, Yi
Kuchaiev, Oleksii
contents Preference datasets are essential for training general-domain, instruction-following language models with Reinforcement Learning from Human Feedback (RLHF). Each subsequent data release raises expectations for future data collection, meaning there is a constant need to advance the quality and diversity of openly available preference data. To address this need, we introduce HelpSteer3-Preference, a permissively licensed (CC-BY-4.0), high-quality, human-annotated preference dataset comprising of over 40,000 samples. These samples span diverse real-world applications of large language models (LLMs), including tasks relating to STEM, coding and multilingual scenarios. Using HelpSteer3-Preference, we train Reward Models (RMs) that achieve top performance on RM-Bench (82.4%) and JudgeBench (73.7%). This represents a substantial improvement (~10% absolute) over the previously best-reported results from existing RMs. We demonstrate HelpSteer3-Preference can also be applied to train Generative RMs and how policy models can be aligned with RLHF using our RMs. Dataset (CC-BY-4.0): https://huggingface.co/datasets/nvidia/HelpSteer3#preference Models (NVIDIA Open Model): https://huggingface.co/collections/nvidia/reward-models-68377c5955575f71fcc7a2a3
format Preprint
id arxiv_https___arxiv_org_abs_2505_11475
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
Wang, Zhilin
Zeng, Jiaqi
Delalleau, Olivier
Shin, Hoo-Chang
Soares, Felipe
Bukharin, Alexander
Evans, Ellie
Dong, Yi
Kuchaiev, Oleksii
Computation and Language
Artificial Intelligence
Machine Learning
Preference datasets are essential for training general-domain, instruction-following language models with Reinforcement Learning from Human Feedback (RLHF). Each subsequent data release raises expectations for future data collection, meaning there is a constant need to advance the quality and diversity of openly available preference data. To address this need, we introduce HelpSteer3-Preference, a permissively licensed (CC-BY-4.0), high-quality, human-annotated preference dataset comprising of over 40,000 samples. These samples span diverse real-world applications of large language models (LLMs), including tasks relating to STEM, coding and multilingual scenarios. Using HelpSteer3-Preference, we train Reward Models (RMs) that achieve top performance on RM-Bench (82.4%) and JudgeBench (73.7%). This represents a substantial improvement (~10% absolute) over the previously best-reported results from existing RMs. We demonstrate HelpSteer3-Preference can also be applied to train Generative RMs and how policy models can be aligned with RLHF using our RMs. Dataset (CC-BY-4.0): https://huggingface.co/datasets/nvidia/HelpSteer3#preference Models (NVIDIA Open Model): https://huggingface.co/collections/nvidia/reward-models-68377c5955575f71fcc7a2a3
title HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.11475