Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Junkang, Xie, Yuexiang, Yang, Zhengyi, Wu, Jiancan, Chen, Jiawei, Gao, Jinyang, Ding, Bolin, Wang, Xiang, He, Xiangnan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910913732280320
author Wu, Junkang
Xie, Yuexiang
Yang, Zhengyi
Wu, Jiancan
Chen, Jiawei
Gao, Jinyang
Ding, Bolin
Wang, Xiang
He, Xiangnan
author_facet Wu, Junkang
Xie, Yuexiang
Yang, Zhengyi
Wu, Jiancan
Chen, Jiawei
Gao, Jinyang
Ding, Bolin
Wang, Xiang
He, Xiangnan
contents This study addresses the challenge of noise in training datasets for Direct Preference Optimization (DPO), a method for aligning Large Language Models (LLMs) with human preferences. We categorize noise into pointwise noise, which includes low-quality data points, and pairwise noise, which encompasses erroneous data pair associations that affect preference rankings. Utilizing Distributionally Robust Optimization (DRO), we enhance DPO's resilience to these types of noise. Our theoretical insights reveal that DPO inherently embeds DRO principles, conferring robustness to pointwise noise, with the regularization coefficient $β$ playing a critical role in its noise resistance. Extending this framework, we introduce Distributionally Robustifying DPO (Dr. DPO), which integrates pairwise robustness by optimizing against worst-case pairwise scenarios. The novel hyperparameter $β'$ in Dr. DPO allows for fine-tuned control over data pair reliability, providing a strategic balance between exploration and exploitation in noisy training environments. Empirical evaluations demonstrate that Dr. DPO substantially improves the quality of generated text and response accuracy in preference datasets, showcasing enhanced performance in both noisy and noise-free settings. The code is available at https://github.com/junkangwu/Dr_DPO.
format Preprint
id arxiv_https___arxiv_org_abs_2407_07880
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
Wu, Junkang
Xie, Yuexiang
Yang, Zhengyi
Wu, Jiancan
Chen, Jiawei
Gao, Jinyang
Ding, Bolin
Wang, Xiang
He, Xiangnan
Machine Learning
Artificial Intelligence
Computation and Language
This study addresses the challenge of noise in training datasets for Direct Preference Optimization (DPO), a method for aligning Large Language Models (LLMs) with human preferences. We categorize noise into pointwise noise, which includes low-quality data points, and pairwise noise, which encompasses erroneous data pair associations that affect preference rankings. Utilizing Distributionally Robust Optimization (DRO), we enhance DPO's resilience to these types of noise. Our theoretical insights reveal that DPO inherently embeds DRO principles, conferring robustness to pointwise noise, with the regularization coefficient $β$ playing a critical role in its noise resistance. Extending this framework, we introduce Distributionally Robustifying DPO (Dr. DPO), which integrates pairwise robustness by optimizing against worst-case pairwise scenarios. The novel hyperparameter $β'$ in Dr. DPO allows for fine-tuned control over data pair reliability, providing a strategic balance between exploration and exploitation in noisy training environments. Empirical evaluations demonstrate that Dr. DPO substantially improves the quality of generated text and response accuracy in preference datasets, showcasing enhanced performance in both noisy and noise-free settings. The code is available at https://github.com/junkangwu/Dr_DPO.
title Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2407.07880