Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rafailov, Rafael, Sharma, Archit, Mitchell, Eric, Ermon, Stefano, Manning, Christopher D., Finn, Chelsea
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!