BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917689041092608 |
|---|---|
| author | Pandey, Gaurav Nandwani, Yatin Naseem, Tahira Mishra, Mayank Xu, Guangxuan Raghu, Dinesh Joshi, Sachindra Munawar, Asim Astudillo, Ramón Fernandez |
| author_facet | Pandey, Gaurav Nandwani, Yatin Naseem, Tahira Mishra, Mayank Xu, Guangxuan Raghu, Dinesh Joshi, Sachindra Munawar, Asim Astudillo, Ramón Fernandez |
| contents | Distribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likelihood Calibration (SLiC), Direct Preference Optimization (DPO) and its variants. We identify high variance of the gradient estimate as the primary reason for the lack of success of these methods and propose a self-normalized baseline to reduce the variance. We further generalize the target distribution in DPG, GDC and DPO by using Bayes' rule to define the reward-conditioned posterior. The resulting approach, referred to as BRAIn - Bayesian Reward-conditioned Amortized Inference acts as a bridge between distribution matching methods and DPO and significantly outperforms prior art in summarization and Antropic HH tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_02479 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback Pandey, Gaurav Nandwani, Yatin Naseem, Tahira Mishra, Mayank Xu, Guangxuan Raghu, Dinesh Joshi, Sachindra Munawar, Asim Astudillo, Ramón Fernandez Machine Learning Artificial Intelligence Computation and Language Human-Computer Interaction Distribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likelihood Calibration (SLiC), Direct Preference Optimization (DPO) and its variants. We identify high variance of the gradient estimate as the primary reason for the lack of success of these methods and propose a self-normalized baseline to reduce the variance. We further generalize the target distribution in DPG, GDC and DPO by using Bayes' rule to define the reward-conditioned posterior. The resulting approach, referred to as BRAIn - Bayesian Reward-conditioned Amortized Inference acts as a bridge between distribution matching methods and DPO and significantly outperforms prior art in summarization and Antropic HH tasks. |
| title | BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback |
| topic | Machine Learning Artificial Intelligence Computation and Language Human-Computer Interaction |
| url | https://arxiv.org/abs/2402.02479 |