BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Kaiwen, Yao, Hongwei, Chen, Yufei, Li, Ziyun, Qiao, Tong, Qin, Zhan, Wang, Cong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912413206446080
author Duan, Kaiwen
Yao, Hongwei
Chen, Yufei
Li, Ziyun
Qiao, Tong
Qin, Zhan
Wang, Cong
author_facet Duan, Kaiwen
Yao, Hongwei
Chen, Yufei
Li, Ziyun
Qiao, Tong
Qin, Zhan
Wang, Cong
contents Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning text-to-image (T2I) models with human preferences. However, RLHF's feedback mechanism also opens new pathways for adversaries. This paper demonstrates the feasibility of hijacking T2I models by poisoning a small fraction of preference data with natural-appearing examples. Specifically, we propose BadReward, a stealthy clean-label poisoning attack targeting the reward model in multi-modal RLHF. BadReward operates by inducing feature collisions between visually contradicted preference data instances, thereby corrupting the reward model and indirectly compromising the T2I model's integrity. Unlike existing alignment poisoning techniques focused on single (text) modality, BadReward is independent of the preference annotation process, enhancing its stealth and practical threat. Extensive experiments on popular T2I models show that BadReward can consistently guide the generation towards improper outputs, such as biased or violent imagery, for targeted concepts. Our findings underscore the amplified threat landscape for RLHF in multi-modal systems, highlighting the urgent need for robust defenses. Disclaimer. This paper contains uncensored toxic content that might be offensive or disturbing to the readers.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03234
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
Duan, Kaiwen
Yao, Hongwei
Chen, Yufei
Li, Ziyun
Qiao, Tong
Qin, Zhan
Wang, Cong
Machine Learning
Artificial Intelligence
Cryptography and Security
Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning text-to-image (T2I) models with human preferences. However, RLHF's feedback mechanism also opens new pathways for adversaries. This paper demonstrates the feasibility of hijacking T2I models by poisoning a small fraction of preference data with natural-appearing examples. Specifically, we propose BadReward, a stealthy clean-label poisoning attack targeting the reward model in multi-modal RLHF. BadReward operates by inducing feature collisions between visually contradicted preference data instances, thereby corrupting the reward model and indirectly compromising the T2I model's integrity. Unlike existing alignment poisoning techniques focused on single (text) modality, BadReward is independent of the preference annotation process, enhancing its stealth and practical threat. Extensive experiments on popular T2I models show that BadReward can consistently guide the generation towards improper outputs, such as biased or violent imagery, for targeted concepts. Our findings underscore the amplified threat landscape for RLHF in multi-modal systems, highlighting the urgent need for robust defenses. Disclaimer. This paper contains uncensored toxic content that might be offensive or disturbing to the readers.
title BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2506.03234