GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kwon, Oh Joon, Matsunaga, Daiki E., Kim, Kee-Eung
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917809892622336
author Kwon, Oh Joon
Matsunaga, Daiki E.
Kim, Kee-Eung
author_facet Kwon, Oh Joon
Matsunaga, Daiki E.
Kim, Kee-Eung
contents A critical component of the current generation of language models is preference alignment, which aims to precisely control the model's behavior to meet human needs and values. The most notable among such methods is Reinforcement Learning with Human Feedback (RLHF) and its offline variant Direct Preference Optimization (DPO), both of which seek to maximize a reward model based on human preferences. In particular, DPO derives reward signals directly from the offline preference data, but in doing so overfits the reward signals and generates suboptimal responses that may contain human biases in the dataset. In this work, we propose a practical application of a diversity-seeking RL algorithm called GFlowNet-DPO (GDPO) in an offline preference alignment setting to curtail such challenges. Empirical results show GDPO can generate far more diverse responses than the baseline methods that are still relatively aligned with human values in dialog generation and summarization tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15096
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets
Kwon, Oh Joon
Matsunaga, Daiki E.
Kim, Kee-Eung
Artificial Intelligence
A critical component of the current generation of language models is preference alignment, which aims to precisely control the model's behavior to meet human needs and values. The most notable among such methods is Reinforcement Learning with Human Feedback (RLHF) and its offline variant Direct Preference Optimization (DPO), both of which seek to maximize a reward model based on human preferences. In particular, DPO derives reward signals directly from the offline preference data, but in doing so overfits the reward signals and generates suboptimal responses that may contain human biases in the dataset. In this work, we propose a practical application of a diversity-seeking RL algorithm called GFlowNet-DPO (GDPO) in an offline preference alignment setting to curtail such challenges. Empirical results show GDPO can generate far more diverse responses than the baseline methods that are still relatively aligned with human values in dialog generation and summarization tasks.
title GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets
topic Artificial Intelligence
url https://arxiv.org/abs/2410.15096