Diverse Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lanchantin, Jack, Chen, Angelica, Dhuliawala, Shehzaad, Yu, Ping, Weston, Jason, Sukhbaatar, Sainbayar, Kulikov, Ilia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913852829990912
author Lanchantin, Jack
Chen, Angelica
Dhuliawala, Shehzaad
Yu, Ping
Weston, Jason
Sukhbaatar, Sainbayar
Kulikov, Ilia
author_facet Lanchantin, Jack
Chen, Angelica
Dhuliawala, Shehzaad
Yu, Ping
Weston, Jason
Sukhbaatar, Sainbayar
Kulikov, Ilia
contents Post-training of language models, either through reinforcement learning, preference optimization or supervised finetuning, tends to sharpen the output probability distribution and reduce the diversity of generated responses. This is particularly a problem for creative generative tasks where varied responses are desired. In this work we introduce Diverse Preference Optimization (DivPO), an optimization method which learns to generate much more diverse responses than standard pipelines, while maintaining the quality of the generations. In DivPO, preference pairs are selected by first considering a pool of responses, and a measure of diversity among them, and selecting chosen examples as being more rare but high quality, while rejected examples are more common, but low quality. DivPO results in generating 45.6% more diverse persona attributes, and a 74.6% increase in story diversity, while maintaining similar win rates as standard baselines. On general instruction following, DivPO results in a 46.2% increase in diversity, and a 2.4% winrate improvement compared to DPO.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18101
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diverse Preference Optimization
Lanchantin, Jack
Chen, Angelica
Dhuliawala, Shehzaad
Yu, Ping
Weston, Jason
Sukhbaatar, Sainbayar
Kulikov, Ilia
Computation and Language
Post-training of language models, either through reinforcement learning, preference optimization or supervised finetuning, tends to sharpen the output probability distribution and reduce the diversity of generated responses. This is particularly a problem for creative generative tasks where varied responses are desired. In this work we introduce Diverse Preference Optimization (DivPO), an optimization method which learns to generate much more diverse responses than standard pipelines, while maintaining the quality of the generations. In DivPO, preference pairs are selected by first considering a pool of responses, and a measure of diversity among them, and selecting chosen examples as being more rare but high quality, while rejected examples are more common, but low quality. DivPO results in generating 45.6% more diverse persona attributes, and a 74.6% increase in story diversity, while maintaining similar win rates as standard baselines. On general instruction following, DivPO results in a 46.2% increase in diversity, and a 2.4% winrate improvement compared to DPO.
title Diverse Preference Optimization
topic Computation and Language
url https://arxiv.org/abs/2501.18101