DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909336831262720 |
|---|---|
| author | Lin, Tzu-Han Li, Chen-An Lee, Hung-yi Chen, Yun-Nung |
| author_facet | Lin, Tzu-Han Li, Chen-An Lee, Hung-yi Chen, Yun-Nung |
| contents | Reinforcement learning from human feedback (RLHF) is a popular strategy for aligning large language models (LLMs) with desired behaviors. Reward modeling is a crucial step in RLHF. However, collecting paired preference data for training reward models is often costly and time-consuming, especially for domain-specific preferences requiring expert annotation. To address this challenge, we propose the \textbf{Do}main knowled\textbf{ge} merged \textbf{R}eward \textbf{M}odel (DogeRM), a novel framework that integrates domain-specific knowledge into a general reward model by model merging. The experiments demonstrate that DogeRM enhances performance across different benchmarks and provide a detailed analysis showcasing the effects of model merging, showing the great potential of facilitating model alignment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_01470 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging Lin, Tzu-Han Li, Chen-An Lee, Hung-yi Chen, Yun-Nung Computation and Language Reinforcement learning from human feedback (RLHF) is a popular strategy for aligning large language models (LLMs) with desired behaviors. Reward modeling is a crucial step in RLHF. However, collecting paired preference data for training reward models is often costly and time-consuming, especially for domain-specific preferences requiring expert annotation. To address this challenge, we propose the \textbf{Do}main knowled\textbf{ge} merged \textbf{R}eward \textbf{M}odel (DogeRM), a novel framework that integrates domain-specific knowledge into a general reward model by model merging. The experiments demonstrate that DogeRM enhances performance across different benchmarks and provide a detailed analysis showcasing the effects of model merging, showing the great potential of facilitating model alignment. |
| title | DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2407.01470 |