Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miao, Yuchun, Ding, Liang, Zhang, Sen, Bao, Rong, Zhang, Lefei, Tao, Dacheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!