Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Chris Yuhao, Zeng, Liang, Liu, Jiacai, Yan, Rui, He, Jujie, Wang, Chaojie, Yan, Shuicheng, Liu, Yang, Zhou, Yahui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910664161755136
author Liu, Chris Yuhao
Zeng, Liang
Liu, Jiacai
Yan, Rui
He, Jujie
Wang, Chaojie
Yan, Shuicheng
Liu, Yang
Zhou, Yahui
author_facet Liu, Chris Yuhao
Zeng, Liang
Liu, Jiacai
Yan, Rui
He, Jujie
Wang, Chaojie
Yan, Shuicheng
Liu, Yang
Zhou, Yahui
contents In this report, we introduce a collection of methods to enhance reward modeling for LLMs, focusing specifically on data-centric techniques. We propose effective data selection and filtering strategies for curating high-quality open-source preference datasets, culminating in the Skywork-Reward data collection, which contains only 80K preference pairs -- significantly smaller than existing datasets. Using this curated dataset, we developed the Skywork-Reward model series -- Skywork-Reward-Gemma-27B and Skywork-Reward-Llama-3.1-8B -- with the former currently holding the top position on the RewardBench leaderboard. Notably, our techniques and datasets have directly enhanced the performance of many top-ranked models on RewardBench, highlighting the practical impact of our contributions in real-world preference learning applications.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18451
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Liu, Chris Yuhao
Zeng, Liang
Liu, Jiacai
Yan, Rui
He, Jujie
Wang, Chaojie
Yan, Shuicheng
Liu, Yang
Zhou, Yahui
Artificial Intelligence
Computation and Language
In this report, we introduce a collection of methods to enhance reward modeling for LLMs, focusing specifically on data-centric techniques. We propose effective data selection and filtering strategies for curating high-quality open-source preference datasets, culminating in the Skywork-Reward data collection, which contains only 80K preference pairs -- significantly smaller than existing datasets. Using this curated dataset, we developed the Skywork-Reward model series -- Skywork-Reward-Gemma-27B and Skywork-Reward-Llama-3.1-8B -- with the former currently holding the top position on the RewardBench leaderboard. Notably, our techniques and datasets have directly enhanced the performance of many top-ranked models on RewardBench, highlighting the practical impact of our contributions in real-world preference learning applications.
title Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.18451