DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Tzu-Han, Li, Chen-An, Lee, Hung-yi, Chen, Yun-Nung
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909336831262720
author Lin, Tzu-Han
Li, Chen-An
Lee, Hung-yi
Chen, Yun-Nung
author_facet Lin, Tzu-Han
Li, Chen-An
Lee, Hung-yi
Chen, Yun-Nung
contents Reinforcement learning from human feedback (RLHF) is a popular strategy for aligning large language models (LLMs) with desired behaviors. Reward modeling is a crucial step in RLHF. However, collecting paired preference data for training reward models is often costly and time-consuming, especially for domain-specific preferences requiring expert annotation. To address this challenge, we propose the \textbf{Do}main knowled\textbf{ge} merged \textbf{R}eward \textbf{M}odel (DogeRM), a novel framework that integrates domain-specific knowledge into a general reward model by model merging. The experiments demonstrate that DogeRM enhances performance across different benchmarks and provide a detailed analysis showcasing the effects of model merging, showing the great potential of facilitating model alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2407_01470
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging
Lin, Tzu-Han
Li, Chen-An
Lee, Hung-yi
Chen, Yun-Nung
Computation and Language
Reinforcement learning from human feedback (RLHF) is a popular strategy for aligning large language models (LLMs) with desired behaviors. Reward modeling is a crucial step in RLHF. However, collecting paired preference data for training reward models is often costly and time-consuming, especially for domain-specific preferences requiring expert annotation. To address this challenge, we propose the \textbf{Do}main knowled\textbf{ge} merged \textbf{R}eward \textbf{M}odel (DogeRM), a novel framework that integrates domain-specific knowledge into a general reward model by model merging. The experiments demonstrate that DogeRM enhances performance across different benchmarks and provide a detailed analysis showcasing the effects of model merging, showing the great potential of facilitating model alignment.
title DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging
topic Computation and Language
url https://arxiv.org/abs/2407.01470