BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Duan, Kaiwen, Yao, Hongwei, Chen, Yufei, Li, Ziyun, Qiao, Tong, Qin, Zhan, Wang, Cong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents
von: Wang, Xuan, et al.
Veröffentlicht: (2025)
von: Wang, Xuan, et al.
Veröffentlicht: (2025)
BadSampler: Harnessing the Power of Catastrophic Forgetting to Poison Byzantine-robust Federated Learning
von: Liu, Yi, et al.
Veröffentlicht: (2024)
von: Liu, Yi, et al.
Veröffentlicht: (2024)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
von: Liu, Tiantian, et al.
Veröffentlicht: (2024)
von: Liu, Tiantian, et al.
Veröffentlicht: (2024)
Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models
von: Shan, Shawn, et al.
Veröffentlicht: (2023)
von: Shan, Shawn, et al.
Veröffentlicht: (2023)
On the Feasibility of Poisoning Text-to-Image AI Models via Adversarial Mislabeling
von: Wu, Stanley, et al.
Veröffentlicht: (2025)
von: Wu, Stanley, et al.
Veröffentlicht: (2025)
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
von: Xu, Yinglun, et al.
Veröffentlicht: (2024)
von: Xu, Yinglun, et al.
Veröffentlicht: (2024)
RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis
von: Wang, Jianwei, et al.
Veröffentlicht: (2025)
von: Wang, Jianwei, et al.
Veröffentlicht: (2025)
SSCL-BW: Sample-Specific Clean-Label Backdoor Watermarking for Dataset Ownership Verification
von: Wang, Yingjia, et al.
Veröffentlicht: (2025)
von: Wang, Yingjia, et al.
Veröffentlicht: (2025)
A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
von: Yao, Yifan, et al.
Veröffentlicht: (2023)
von: Yao, Yifan, et al.
Veröffentlicht: (2023)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2023)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2023)
A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and Triggers
von: Wu, Zhixiao, et al.
Veröffentlicht: (2025)
von: Wu, Zhixiao, et al.
Veröffentlicht: (2025)
Clean-Label Physical Backdoor Attacks with Data Distillation
von: Dao, Thinh, et al.
Veröffentlicht: (2024)
von: Dao, Thinh, et al.
Veröffentlicht: (2024)
The Stronger the Diffusion Model, the Easier the Backdoor: Data Poisoning to Induce Copyright Breaches Without Adjusting Finetuning Pipeline
von: Wang, Haonan, et al.
Veröffentlicht: (2024)
von: Wang, Haonan, et al.
Veröffentlicht: (2024)
Strategic Sample Selection for Improved Clean-Label Backdoor Attacks in Text Classification
von: Kirci, Onur Alp, et al.
Veröffentlicht: (2025)
von: Kirci, Onur Alp, et al.
Veröffentlicht: (2025)
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2024)
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2024)
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
Selection-Based Vulnerabilities: Clean-Label Backdoor Attacks in Active Learning
von: Zhi, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhi, Yuhan, et al.
Veröffentlicht: (2025)
Exposing Vulnerabilities in RL: A Novel Stealthy Backdoor Attack through Reward Poisoning
von: Zhang, Bokang, et al.
Veröffentlicht: (2025)
von: Zhang, Bokang, et al.
Veröffentlicht: (2025)
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
von: Liu, Yule, et al.
Veröffentlicht: (2025)
von: Liu, Yule, et al.
Veröffentlicht: (2025)
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
von: Chen, Depeng, et al.
Veröffentlicht: (2024)
von: Chen, Depeng, et al.
Veröffentlicht: (2024)
FIT-Print: Towards False-claim-resistant Model Ownership Verification via Targeted Fingerprint
von: Shao, Shuo, et al.
Veröffentlicht: (2025)
von: Shao, Shuo, et al.
Veröffentlicht: (2025)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
von: Shao, Shuo, et al.
Veröffentlicht: (2024)
von: Shao, Shuo, et al.
Veröffentlicht: (2024)
Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
von: Sasnauskas, Paulius, et al.
Veröffentlicht: (2025)
von: Sasnauskas, Paulius, et al.
Veröffentlicht: (2025)
Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation
von: Shao, Shuo, et al.
Veröffentlicht: (2025)
von: Shao, Shuo, et al.
Veröffentlicht: (2025)
DFB: A Data-Free, Low-Budget, and High-Efficacy Clean-Label Backdoor Attack
von: Ma, Binhao, et al.
Veröffentlicht: (2023)
von: Ma, Binhao, et al.
Veröffentlicht: (2023)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
Towards Sample-specific Backdoor Attack with Clean Labels via Attribute Trigger
von: Zhu, Mingyan, et al.
Veröffentlicht: (2023)
von: Zhu, Mingyan, et al.
Veröffentlicht: (2023)
ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
SoK: Large Language Model Copyright Auditing via Fingerprinting
von: Shao, Shuo, et al.
Veröffentlicht: (2025)
von: Shao, Shuo, et al.
Veröffentlicht: (2025)
Post-Training Local LLM Agents for Linux Privilege Escalation with Verifiable Rewards
von: Normann, Philipp, et al.
Veröffentlicht: (2026)
von: Normann, Philipp, et al.
Veröffentlicht: (2026)
A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming
von: Ding, Yizhong
Veröffentlicht: (2025)
von: Ding, Yizhong
Veröffentlicht: (2025)
MCP-ITP: An Automated Framework for Implicit Tool Poisoning in MCP
von: Li, Ruiqi, et al.
Veröffentlicht: (2026)
von: Li, Ruiqi, et al.
Veröffentlicht: (2026)
Poisoned Acoustics
von: Dahme, Harrison
Veröffentlicht: (2026)
von: Dahme, Harrison
Veröffentlicht: (2026)
Less is more? Rewards in RL for Cyber Defence
von: Bates, Elizabeth, et al.
Veröffentlicht: (2025)
von: Bates, Elizabeth, et al.
Veröffentlicht: (2025)
ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
von: Liang, Jiacheng, et al.
Veröffentlicht: (2026)
von: Liang, Jiacheng, et al.
Veröffentlicht: (2026)
Knowledge Poisoning Attacks on Medical Multi-Modal Retrieval-Augmented Generation
von: Yang, Peiru, et al.
Veröffentlicht: (2026)
von: Yang, Peiru, et al.
Veröffentlicht: (2026)
ControlNET: A Firewall for RAG-based LLM System
von: Yao, Hongwei, et al.
Veröffentlicht: (2025)
von: Yao, Hongwei, et al.
Veröffentlicht: (2025)
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
von: Chen, Tianxin, et al.
Veröffentlicht: (2026)
von: Chen, Tianxin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents
von: Wang, Xuan, et al.
Veröffentlicht: (2025) -
BadSampler: Harnessing the Power of Catastrophic Forgetting to Poison Byzantine-robust Federated Learning
von: Liu, Yi, et al.
Veröffentlicht: (2024) -
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
von: Tie, Guiyao, et al.
Veröffentlicht: (2026) -
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
von: Liu, Tiantian, et al.
Veröffentlicht: (2024) -
Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models
von: Shan, Shawn, et al.
Veröffentlicht: (2023)