Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yinglun, Gumaste, Rohan, Singh, Gagandeep |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Binary Reward Labeling: Bridging Offline Preference and Reward-Based Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024)
by: Xu, Yinglun, et al.
Published: (2024)
GFCL: A GRU-based Federated Continual Learning Framework against Data Poisoning Attacks in IoV
by: Talpur, Anum, et al.
Published: (2022)
by: Talpur, Anum, et al.
Published: (2022)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
by: Vega, Jason, et al.
Published: (2023)
by: Vega, Jason, et al.
Published: (2023)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
SurvAttack: Black-Box Attack On Survival Models through Ontology-Informed EHR Perturbation
by: Kerdabadi, Mohsen Nayebi, et al.
Published: (2024)
by: Kerdabadi, Mohsen Nayebi, et al.
Published: (2024)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
by: Duan, Kaiwen, et al.
Published: (2025)
by: Duan, Kaiwen, et al.
Published: (2025)
BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets
by: Gong, Chen, et al.
Published: (2022)
by: Gong, Chen, et al.
Published: (2022)
Counter-Samples: A Stateless Strategy to Neutralize Black Box Adversarial Attacks
by: Bokobza, Roey, et al.
Published: (2024)
by: Bokobza, Roey, et al.
Published: (2024)
Have You Poisoned My Data? Defending Neural Networks against Data Poisoning
by: De Gaspari, Fabio, et al.
Published: (2024)
by: De Gaspari, Fabio, et al.
Published: (2024)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
by: Carnerero-Cano, Javier, et al.
Published: (2026)
by: Carnerero-Cano, Javier, et al.
Published: (2026)
Prompt-Induced Over-Generation as Denial-of-Service: A Black-Box Attack-Side Benchmark
by: Manu, et al.
Published: (2025)
by: Manu, et al.
Published: (2025)
SAFELOC: Overcoming Data Poisoning Attacks in Heterogeneous Federated Machine Learning for Indoor Localization
by: Singampalli, Akhil, et al.
Published: (2024)
by: Singampalli, Akhil, et al.
Published: (2024)
EAB-FL: Exacerbating Algorithmic Bias through Model Poisoning Attacks in Federated Learning
by: Meerza, Syed Irfan Ali, et al.
Published: (2024)
by: Meerza, Syed Irfan Ali, et al.
Published: (2024)
Turning Black Box into White Box: Dataset Distillation Leaks
by: Chen, Huajie, et al.
Published: (2026)
by: Chen, Huajie, et al.
Published: (2026)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
by: Sasnauskas, Paulius, et al.
Published: (2025)
by: Sasnauskas, Paulius, et al.
Published: (2025)
UNIDOOR: A Universal Framework for Action-Level Backdoor Attacks in Deep Reinforcement Learning
by: Ma, Oubo, et al.
Published: (2025)
by: Ma, Oubo, et al.
Published: (2025)
Semantic Chameleon: Corpus-Dependent Poisoning Attacks and Defenses in RAG Systems
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
BLAST: A Stealthy Backdoor Leverage Attack against Cooperative Multi-Agent Deep Reinforcement Learning based Systems
by: Fang, Jing, et al.
Published: (2025)
by: Fang, Jing, et al.
Published: (2025)
Protecting Federated Learning from Extreme Model Poisoning Attacks via Multidimensional Time Series Anomaly Detection
by: Gabrielli, Edoardo, et al.
Published: (2023)
by: Gabrielli, Edoardo, et al.
Published: (2023)
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
by: Vyas, Sanyam, et al.
Published: (2025)
by: Vyas, Sanyam, et al.
Published: (2025)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025)
by: Foerster, Hanna, et al.
Published: (2025)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
by: Di, Jimmy Z., et al.
Published: (2022)
by: Di, Jimmy Z., et al.
Published: (2022)
Client-Side Patching against Backdoor Attacks in Federated Learning
by: Molina-Coronado, Borja
Published: (2024)
by: Molina-Coronado, Borja
Published: (2024)
ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking
by: Li, Yunzhe, et al.
Published: (2025)
by: Li, Yunzhe, et al.
Published: (2025)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
by: Sitawarin, Chawin, et al.
Published: (2024)
by: Sitawarin, Chawin, et al.
Published: (2024)
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
by: Domico, Kyle, et al.
Published: (2025)
by: Domico, Kyle, et al.
Published: (2025)
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
by: Halloran, John T., et al.
Published: (2026)
by: Halloran, John T., et al.
Published: (2026)
A White-Box Adversarial Attack Against a Digital Twin
by: Patterson, Wilson, et al.
Published: (2022)
by: Patterson, Wilson, et al.
Published: (2022)
Defending the Edge: Representative-Attention Defense against Backdoor Attacks in Federated Learning
by: Obioma, Chibueze Peace, et al.
Published: (2025)
by: Obioma, Chibueze Peace, et al.
Published: (2025)
Similarity-based Label Inference Attack against Training and Inference of Split Learning
by: Liu, Junlin, et al.
Published: (2022)
by: Liu, Junlin, et al.
Published: (2022)
Knowledge-Informed Auto-Penetration Testing Based on Reinforcement Learning with Reward Machine
by: Li, Yuanliang, et al.
Published: (2024)
by: Li, Yuanliang, et al.
Published: (2024)
Tricking Retrievers with Influential Tokens: An Efficient Black-Box Corpus Poisoning Attack
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
Cross-Input Certified Training for Universal Perturbations
by: Xu, Changming, et al.
Published: (2024)
by: Xu, Changming, et al.
Published: (2024)
POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
by: Li, Xinyu, et al.
Published: (2025)
by: Li, Xinyu, et al.
Published: (2025)
Leveraging Reinforcement Learning in Red Teaming for Advanced Ransomware Attack Simulations
by: Wang, Cheng, et al.
Published: (2024)
by: Wang, Cheng, et al.
Published: (2024)
Online Poisoning Attack Against Reinforcement Learning under Black-box Environments
by: Li, Jianhui, et al.
Published: (2024)
by: Li, Jianhui, et al.
Published: (2024)
Similar Items
-
Binary Reward Labeling: Bridging Offline Preference and Reward-Based Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024) -
GFCL: A GRU-based Federated Continual Learning Framework against Data Poisoning Attacks in IoV
by: Talpur, Anum, et al.
Published: (2022) -
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
by: Lobo, Elita, et al.
Published: (2024) -
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
by: Vega, Jason, et al.
Published: (2023) -
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)