On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Jiancong, Li, Ziniu, Xie, Xingyu, Getzen, Emily, Fang, Cong, Long, Qi, Su, Weijie J.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!