RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913712883892224 |
|---|---|
| author | Zhao, Hanyang Winata, Genta Indra Das, Anirban Zhang, Shi-Xiong Yao, David D. Tang, Wenpin Sahu, Sambit |
| author_facet | Zhao, Hanyang Winata, Genta Indra Das, Anirban Zhang, Shi-Xiong Yao, David D. Tang, Wenpin Sahu, Sambit |
| contents | Recently, numerous preference optimization algorithms have been introduced as extensions to the Direct Preference Optimization (DPO) family. While these methods have successfully aligned models with human preferences, there is a lack of understanding regarding the contributions of their additional components. Moreover, fair and consistent comparisons are scarce, making it difficult to discern which components genuinely enhance downstream performance. In this work, we propose RainbowPO, a unified framework that demystifies the effectiveness of existing DPO methods by categorizing their key components into seven broad directions. We integrate these components into a single cohesive objective, enhancing the performance of each individual element. Through extensive experiments, we demonstrate that RainbowPO outperforms existing DPO variants. Additionally, we provide insights to guide researchers in developing new DPO methods and assist practitioners in their implementations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_04203 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization Zhao, Hanyang Winata, Genta Indra Das, Anirban Zhang, Shi-Xiong Yao, David D. Tang, Wenpin Sahu, Sambit Artificial Intelligence Recently, numerous preference optimization algorithms have been introduced as extensions to the Direct Preference Optimization (DPO) family. While these methods have successfully aligned models with human preferences, there is a lack of understanding regarding the contributions of their additional components. Moreover, fair and consistent comparisons are scarce, making it difficult to discern which components genuinely enhance downstream performance. In this work, we propose RainbowPO, a unified framework that demystifies the effectiveness of existing DPO methods by categorizing their key components into seven broad directions. We integrate these components into a single cohesive objective, enhancing the performance of each individual element. Through extensive experiments, we demonstrate that RainbowPO outperforms existing DPO variants. Additionally, we provide insights to guide researchers in developing new DPO methods and assist practitioners in their implementations. |
| title | RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2410.04203 |