Group Robust Preference Optimization in Reward-free RLHF

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ramesh, Shyam Sundhar, Hu, Yifan, Chaimalas, Iason, Mehta, Viraj, Sessa, Pier Giuseppe, Ammar, Haitham Bou, Bogunovic, Ilija
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914816710410240
author Ramesh, Shyam Sundhar
Hu, Yifan
Chaimalas, Iason
Mehta, Viraj
Sessa, Pier Giuseppe
Ammar, Haitham Bou
Bogunovic, Ilija
author_facet Ramesh, Shyam Sundhar
Hu, Yifan
Chaimalas, Iason
Mehta, Viraj
Sessa, Pier Giuseppe
Ammar, Haitham Bou
Bogunovic, Ilija
contents Adapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data. While these data often come from diverse labelers' groups (e.g., different demographics, ethnicities, company teams, etc.), traditional RLHF approaches adopt a "one-size-fits-all" approach, i.e., they indiscriminately assume and optimize a single preference model, thus not being robust to unique characteristics and needs of the various groups. To address this limitation, we propose a novel Group Robust Preference Optimization (GRPO) method to align LLMs to individual groups' preferences robustly. Our approach builds upon reward-free direct preference optimization methods, but unlike previous approaches, it seeks a robust policy which maximizes the worst-case group performance. To achieve this, GRPO adaptively and sequentially weights the importance of different groups, prioritizing groups with worse cumulative loss. We theoretically study the feasibility of GRPO and analyze its convergence for the log-linear policy class. By fine-tuning LLMs with GRPO using diverse group-based global opinion data, we significantly improved performance for the worst-performing groups, reduced loss imbalances across groups, and improved probability accuracies compared to non-robust baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20304
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Group Robust Preference Optimization in Reward-free RLHF
Ramesh, Shyam Sundhar
Hu, Yifan
Chaimalas, Iason
Mehta, Viraj
Sessa, Pier Giuseppe
Ammar, Haitham Bou
Bogunovic, Ilija
Computation and Language
Machine Learning
Adapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data. While these data often come from diverse labelers' groups (e.g., different demographics, ethnicities, company teams, etc.), traditional RLHF approaches adopt a "one-size-fits-all" approach, i.e., they indiscriminately assume and optimize a single preference model, thus not being robust to unique characteristics and needs of the various groups. To address this limitation, we propose a novel Group Robust Preference Optimization (GRPO) method to align LLMs to individual groups' preferences robustly. Our approach builds upon reward-free direct preference optimization methods, but unlike previous approaches, it seeks a robust policy which maximizes the worst-case group performance. To achieve this, GRPO adaptively and sequentially weights the importance of different groups, prioritizing groups with worse cumulative loss. We theoretically study the feasibility of GRPO and analyze its convergence for the log-linear policy class. By fine-tuning LLMs with GRPO using diverse group-based global opinion data, we significantly improved performance for the worst-performing groups, reduced loss imbalances across groups, and improved probability accuracies compared to non-robust baselines.
title Group Robust Preference Optimization in Reward-free RLHF
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2405.20304