RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Hanyang, Winata, Genta Indra, Das, Anirban, Zhang, Shi-Xiong, Yao, David D., Tang, Wenpin, Sahu, Sambit
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913712883892224
author Zhao, Hanyang
Winata, Genta Indra
Das, Anirban
Zhang, Shi-Xiong
Yao, David D.
Tang, Wenpin
Sahu, Sambit
author_facet Zhao, Hanyang
Winata, Genta Indra
Das, Anirban
Zhang, Shi-Xiong
Yao, David D.
Tang, Wenpin
Sahu, Sambit
contents Recently, numerous preference optimization algorithms have been introduced as extensions to the Direct Preference Optimization (DPO) family. While these methods have successfully aligned models with human preferences, there is a lack of understanding regarding the contributions of their additional components. Moreover, fair and consistent comparisons are scarce, making it difficult to discern which components genuinely enhance downstream performance. In this work, we propose RainbowPO, a unified framework that demystifies the effectiveness of existing DPO methods by categorizing their key components into seven broad directions. We integrate these components into a single cohesive objective, enhancing the performance of each individual element. Through extensive experiments, we demonstrate that RainbowPO outperforms existing DPO variants. Additionally, we provide insights to guide researchers in developing new DPO methods and assist practitioners in their implementations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04203
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization
Zhao, Hanyang
Winata, Genta Indra
Das, Anirban
Zhang, Shi-Xiong
Yao, David D.
Tang, Wenpin
Sahu, Sambit
Artificial Intelligence
Recently, numerous preference optimization algorithms have been introduced as extensions to the Direct Preference Optimization (DPO) family. While these methods have successfully aligned models with human preferences, there is a lack of understanding regarding the contributions of their additional components. Moreover, fair and consistent comparisons are scarce, making it difficult to discern which components genuinely enhance downstream performance. In this work, we propose RainbowPO, a unified framework that demystifies the effectiveness of existing DPO methods by categorizing their key components into seven broad directions. We integrate these components into a single cohesive objective, enhancing the performance of each individual element. Through extensive experiments, we demonstrate that RainbowPO outperforms existing DPO variants. Additionally, we provide insights to guide researchers in developing new DPO methods and assist practitioners in their implementations.
title RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization
topic Artificial Intelligence
url https://arxiv.org/abs/2410.04203