When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Escamilla, Jose Efraim Aguilar, Hong, Haoyang, Li, Jiawei, Zhao, Haoyu, Zhang, Xuezhou, Hong, Sanghyun, Wang, Huazheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hard Work Does Not Always Pay Off: Poisoning Attacks on Neural Architecture Search
von: Coalson, Zachary, et al.
Veröffentlicht: (2024)
von: Coalson, Zachary, et al.
Veröffentlicht: (2024)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
von: Sasnauskas, Paulius, et al.
Veröffentlicht: (2025)
von: Sasnauskas, Paulius, et al.
Veröffentlicht: (2025)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025)
Efficient Reinforcement Learning in Probabilistic Reward Machines
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2024)
Preference Poisoning Attacks on Reward Model Learning
von: Wu, Junlin, et al.
Veröffentlicht: (2024)
von: Wu, Junlin, et al.
Veröffentlicht: (2024)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
Certified Robustness to Clean-Label Poisoning Using Diffusion Denoising
von: Hong, Sanghyun, et al.
Veröffentlicht: (2024)
von: Hong, Sanghyun, et al.
Veröffentlicht: (2024)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
von: Duan, Kaiwen, et al.
Veröffentlicht: (2025)
von: Duan, Kaiwen, et al.
Veröffentlicht: (2025)
Robust Thompson Sampling Algorithms Against Reward Poisoning Attacks
von: Xu, Yinglun, et al.
Veröffentlicht: (2024)
von: Xu, Yinglun, et al.
Veröffentlicht: (2024)
RA-PbRL: Provably Efficient Risk-Aware Preference-Based Reinforcement Learning
von: Zhao, Yujie, et al.
Veröffentlicht: (2024)
von: Zhao, Yujie, et al.
Veröffentlicht: (2024)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
von: Zhang, Junkai, et al.
Veröffentlicht: (2023)
von: Zhang, Junkai, et al.
Veröffentlicht: (2023)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
von: Wen, Yuxin, et al.
Veröffentlicht: (2024)
von: Wen, Yuxin, et al.
Veröffentlicht: (2024)
Towards Poisoning Fair Representations
von: Liu, Tianci, et al.
Veröffentlicht: (2023)
von: Liu, Tianci, et al.
Veröffentlicht: (2023)
What Can You Do When You Have Zero Rewards During RL?
von: Prakash, Jatin, et al.
Veröffentlicht: (2025)
von: Prakash, Jatin, et al.
Veröffentlicht: (2025)
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
von: Xu, Yinglun, et al.
Veröffentlicht: (2024)
von: Xu, Yinglun, et al.
Veröffentlicht: (2024)
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
von: Xie, Tong, et al.
Veröffentlicht: (2025)
von: Xie, Tong, et al.
Veröffentlicht: (2025)
Beyond Static Bias: Adaptive Multi-Fidelity Bandits with Improving Proxies
von: Lu, Muyun, et al.
Veröffentlicht: (2026)
von: Lu, Muyun, et al.
Veröffentlicht: (2026)
Have You Poisoned My Data? Defending Neural Networks against Data Poisoning
von: De Gaspari, Fabio, et al.
Veröffentlicht: (2024)
von: De Gaspari, Fabio, et al.
Veröffentlicht: (2024)
On Robustness of Linear Classifiers to Targeted Data Poisoning
von: Gupta, Nakshatra, et al.
Veröffentlicht: (2025)
von: Gupta, Nakshatra, et al.
Veröffentlicht: (2025)
When Do "More Contexts" Help with Sarcasm Recognition?
von: Nimase, Ojas, et al.
Veröffentlicht: (2024)
von: Nimase, Ojas, et al.
Veröffentlicht: (2024)
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
von: Wei, Yukuan, et al.
Veröffentlicht: (2025)
von: Wei, Yukuan, et al.
Veröffentlicht: (2025)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2025)
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2025)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
von: Boone, Victor, et al.
Veröffentlicht: (2024)
von: Boone, Victor, et al.
Veröffentlicht: (2024)
A Linear Approach to Data Poisoning
von: Flynn, Donald, et al.
Veröffentlicht: (2025)
von: Flynn, Donald, et al.
Veröffentlicht: (2025)
Design-Based Bandits Under Network Interference: Trade-Off Between Regret and Statistical Inference
von: Wang, Zichen, et al.
Veröffentlicht: (2025)
von: Wang, Zichen, et al.
Veröffentlicht: (2025)
Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
Two-Timescale Critic-Actor for Average Reward MDPs with Function Approximation
von: Panda, Prashansa, et al.
Veröffentlicht: (2024)
von: Panda, Prashansa, et al.
Veröffentlicht: (2024)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
Confundo: Learning to Generate Robust Poison for Practical RAG Systems
von: Hu, Haoyang, et al.
Veröffentlicht: (2026)
von: Hu, Haoyang, et al.
Veröffentlicht: (2026)
Sample Complexity Characterization for Linear Contextual MDPs
von: Deng, Junze, et al.
Veröffentlicht: (2024)
von: Deng, Junze, et al.
Veröffentlicht: (2024)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach
von: Ganesh, Swetha, et al.
Veröffentlicht: (2025)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2025)
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
von: Souly, Alexandra, et al.
Veröffentlicht: (2025)
von: Souly, Alexandra, et al.
Veröffentlicht: (2025)
FuncPoison: Poisoning Function Library to Hijack Multi-agent Autonomous Driving Systems
von: Long, Yuzhen, et al.
Veröffentlicht: (2025)
von: Long, Yuzhen, et al.
Veröffentlicht: (2025)
Span-Based Optimal Sample Complexity for Average Reward MDPs
von: Zurek, Matthew, et al.
Veröffentlicht: (2023)
von: Zurek, Matthew, et al.
Veröffentlicht: (2023)
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
von: Huang, Yu, et al.
Veröffentlicht: (2026)
von: Huang, Yu, et al.
Veröffentlicht: (2026)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
Associative Poisoning to Generative Machine Learning
von: Mohus, Mathias Lundteigen, et al.
Veröffentlicht: (2025)
von: Mohus, Mathias Lundteigen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hard Work Does Not Always Pay Off: Poisoning Attacks on Neural Architecture Search
von: Coalson, Zachary, et al.
Veröffentlicht: (2024) -
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024) -
Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
von: Sasnauskas, Paulius, et al.
Veröffentlicht: (2025) -
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025) -
Efficient Reinforcement Learning in Probabilistic Reward Machines
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2024)