No Preference Left Behind: Group Distributional Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Binwei, Cai, Zefan, Chuang, Yun-Shiuan, Yang, Shanglin, Jiang, Ming, Yang, Diyi, Hu, Junjie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913833313894400
author Yao, Binwei
Cai, Zefan
Chuang, Yun-Shiuan
Yang, Shanglin
Jiang, Ming
Yang, Diyi
Hu, Junjie
author_facet Yao, Binwei
Cai, Zefan
Chuang, Yun-Shiuan
Yang, Shanglin
Jiang, Ming
Yang, Diyi
Hu, Junjie
contents Preferences within a group of people are not uniform but follow a distribution. While existing alignment methods like Direct Preference Optimization (DPO) attempt to steer models to reflect human preferences, they struggle to capture the distributional pluralistic preferences within a group. These methods often skew toward dominant preferences, overlooking the diversity of opinions, especially when conflicting preferences arise. To address this issue, we propose Group Distributional Preference Optimization (GDPO), a novel framework that aligns language models with the distribution of preferences within a group by incorporating the concept of beliefs that shape individual preferences. GDPO calibrates a language model using statistical estimation of the group's belief distribution and aligns the model with belief-conditioned preferences, offering a more inclusive alignment framework than traditional methods. In experiments using both synthetic controllable opinion generation and real-world movie review datasets, we show that DPO fails to align with the targeted belief distributions, while GDPO consistently reduces this alignment gap during training. Moreover, our evaluation metrics demonstrate that GDPO outperforms existing approaches in aligning with group distributional preferences, marking a significant advance in pluralistic alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20299
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle No Preference Left Behind: Group Distributional Preference Optimization
Yao, Binwei
Cai, Zefan
Chuang, Yun-Shiuan
Yang, Shanglin
Jiang, Ming
Yang, Diyi
Hu, Junjie
Computation and Language
Preferences within a group of people are not uniform but follow a distribution. While existing alignment methods like Direct Preference Optimization (DPO) attempt to steer models to reflect human preferences, they struggle to capture the distributional pluralistic preferences within a group. These methods often skew toward dominant preferences, overlooking the diversity of opinions, especially when conflicting preferences arise. To address this issue, we propose Group Distributional Preference Optimization (GDPO), a novel framework that aligns language models with the distribution of preferences within a group by incorporating the concept of beliefs that shape individual preferences. GDPO calibrates a language model using statistical estimation of the group's belief distribution and aligns the model with belief-conditioned preferences, offering a more inclusive alignment framework than traditional methods. In experiments using both synthetic controllable opinion generation and real-world movie review datasets, we show that DPO fails to align with the targeted belief distributions, while GDPO consistently reduces this alignment gap during training. Moreover, our evaluation metrics demonstrate that GDPO outperforms existing approaches in aligning with group distributional preferences, marking a significant advance in pluralistic alignment.
title No Preference Left Behind: Group Distributional Preference Optimization
topic Computation and Language
url https://arxiv.org/abs/2412.20299