BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pandey, Gaurav, Nandwani, Yatin, Naseem, Tahira, Mishra, Mayank, Xu, Guangxuan, Raghu, Dinesh, Joshi, Sachindra, Munawar, Asim, Astudillo, Ramón Fernandez
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917689041092608
author Pandey, Gaurav
Nandwani, Yatin
Naseem, Tahira
Mishra, Mayank
Xu, Guangxuan
Raghu, Dinesh
Joshi, Sachindra
Munawar, Asim
Astudillo, Ramón Fernandez
author_facet Pandey, Gaurav
Nandwani, Yatin
Naseem, Tahira
Mishra, Mayank
Xu, Guangxuan
Raghu, Dinesh
Joshi, Sachindra
Munawar, Asim
Astudillo, Ramón Fernandez
contents Distribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likelihood Calibration (SLiC), Direct Preference Optimization (DPO) and its variants. We identify high variance of the gradient estimate as the primary reason for the lack of success of these methods and propose a self-normalized baseline to reduce the variance. We further generalize the target distribution in DPG, GDC and DPO by using Bayes' rule to define the reward-conditioned posterior. The resulting approach, referred to as BRAIn - Bayesian Reward-conditioned Amortized Inference acts as a bridge between distribution matching methods and DPO and significantly outperforms prior art in summarization and Antropic HH tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02479
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
Pandey, Gaurav
Nandwani, Yatin
Naseem, Tahira
Mishra, Mayank
Xu, Guangxuan
Raghu, Dinesh
Joshi, Sachindra
Munawar, Asim
Astudillo, Ramón Fernandez
Machine Learning
Artificial Intelligence
Computation and Language
Human-Computer Interaction
Distribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likelihood Calibration (SLiC), Direct Preference Optimization (DPO) and its variants. We identify high variance of the gradient estimate as the primary reason for the lack of success of these methods and propose a self-normalized baseline to reduce the variance. We further generalize the target distribution in DPG, GDC and DPO by using Bayes' rule to define the reward-conditioned posterior. The resulting approach, referred to as BRAIn - Bayesian Reward-conditioned Amortized Inference acts as a bridge between distribution matching methods and DPO and significantly outperforms prior art in summarization and Antropic HH tasks.
title BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
topic Machine Learning
Artificial Intelligence
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2402.02479