RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Harrison, Phatale, Samrat, Mansoor, Hassan, Mesnard, Thomas, Ferret, Johan, Lu, Kellie, Bishop, Colton, Hall, Ethan, Carbune, Victor, Rastogi, Abhinav, Prakash, Sushant
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!