On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lin, Yong, Seto, Skyler, ter Hoeve, Maartje, Metcalf, Katherine, Theobald, Barry-John, Wang, Xuan, Zhang, Yizhe, Huang, Chen, Zhang, Tong
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!