BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Kaiwen, Yao, Hongwei, Chen, Yufei, Li, Ziyun, Qiao, Tong, Qin, Zhan, Wang, Cong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!