Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nika, Andi, Mandal, Debmalya, Kamalaruban, Parameswaran, Tzannetos, Georgios, Radanović, Goran, Singla, Adish
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!