Online Bandit Learning with Offline Preference Data for Improved RLHF

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Agnihotri, Akhil, Jain, Rahul, Ramachandran, Deepak, Wen, Zheng
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913840114958336
author Agnihotri, Akhil
Jain, Rahul
Ramachandran, Deepak
Wen, Zheng
author_facet Agnihotri, Akhil
Jain, Rahul
Ramachandran, Deepak
Wen, Zheng
contents Reinforcement Learning with Human Feedback (RLHF) is at the core of fine-tuning methods for generative AI models for language and images. Such feedback is often sought as rank or preference feedback from human raters, as opposed to eliciting scores since the latter tends to be noisy. On the other hand, RL theory and algorithms predominantly assume that a reward feedback is available. In particular, approaches for online learning that can be helpful in adaptive data collection via active learning cannot incorporate offline preference data. In this paper, we adopt a finite-armed linear bandit model as a prototypical model of online learning. We consider an offline preference dataset to be available generated by an expert of unknown 'competence'. We propose warmPref-PS, a posterior sampling algorithm for online learning that can be warm-started with an offline dataset with noisy preference feedback. We show that by modeling the 'competence' of the expert that generated it, we are able to use such a dataset most effectively. We support our claims with novel theoretical analysis of its Bayesian regret, as well as, extensive empirical evaluation of an approximate loss function that optimizes for infinitely many arms, and performs substantially better than baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09574
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Online Bandit Learning with Offline Preference Data for Improved RLHF
Agnihotri, Akhil
Jain, Rahul
Ramachandran, Deepak
Wen, Zheng
Machine Learning
Reinforcement Learning with Human Feedback (RLHF) is at the core of fine-tuning methods for generative AI models for language and images. Such feedback is often sought as rank or preference feedback from human raters, as opposed to eliciting scores since the latter tends to be noisy. On the other hand, RL theory and algorithms predominantly assume that a reward feedback is available. In particular, approaches for online learning that can be helpful in adaptive data collection via active learning cannot incorporate offline preference data. In this paper, we adopt a finite-armed linear bandit model as a prototypical model of online learning. We consider an offline preference dataset to be available generated by an expert of unknown 'competence'. We propose warmPref-PS, a posterior sampling algorithm for online learning that can be warm-started with an offline dataset with noisy preference feedback. We show that by modeling the 'competence' of the expert that generated it, we are able to use such a dataset most effectively. We support our claims with novel theoretical analysis of its Bayesian regret, as well as, extensive empirical evaluation of an approximate loss function that optimizes for infinitely many arms, and performs substantially better than baselines.
title Online Bandit Learning with Offline Preference Data for Improved RLHF
topic Machine Learning
url https://arxiv.org/abs/2406.09574