When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Escamilla, Jose Efraim Aguilar, Hong, Haoyang, Li, Jiawei, Zhao, Haoyu, Zhang, Xuezhou, Hong, Sanghyun, Wang, Huazheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hard Work Does Not Always Pay Off: Poisoning Attacks on Neural Architecture Search
by: Coalson, Zachary, et al.
Published: (2024)
by: Coalson, Zachary, et al.
Published: (2024)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
by: Sasnauskas, Paulius, et al.
Published: (2025)
by: Sasnauskas, Paulius, et al.
Published: (2025)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2025)
by: Hong, Kihyuk, et al.
Published: (2025)
Efficient Reinforcement Learning in Probabilistic Reward Machines
by: Lin, Xiaofeng, et al.
Published: (2024)
by: Lin, Xiaofeng, et al.
Published: (2024)
Preference Poisoning Attacks on Reward Model Learning
by: Wu, Junlin, et al.
Published: (2024)
by: Wu, Junlin, et al.
Published: (2024)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
Certified Robustness to Clean-Label Poisoning Using Diffusion Denoising
by: Hong, Sanghyun, et al.
Published: (2024)
by: Hong, Sanghyun, et al.
Published: (2024)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
by: Duan, Kaiwen, et al.
Published: (2025)
by: Duan, Kaiwen, et al.
Published: (2025)
Robust Thompson Sampling Algorithms Against Reward Poisoning Attacks
by: Xu, Yinglun, et al.
Published: (2024)
by: Xu, Yinglun, et al.
Published: (2024)
RA-PbRL: Provably Efficient Risk-Aware Preference-Based Reinforcement Learning
by: Zhao, Yujie, et al.
Published: (2024)
by: Zhao, Yujie, et al.
Published: (2024)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
by: Zhang, Junkai, et al.
Published: (2023)
by: Zhang, Junkai, et al.
Published: (2023)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
Towards Poisoning Fair Representations
by: Liu, Tianci, et al.
Published: (2023)
by: Liu, Tianci, et al.
Published: (2023)
What Can You Do When You Have Zero Rewards During RL?
by: Prakash, Jatin, et al.
Published: (2025)
by: Prakash, Jatin, et al.
Published: (2025)
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024)
by: Xu, Yinglun, et al.
Published: (2024)
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
by: Xie, Tong, et al.
Published: (2025)
by: Xie, Tong, et al.
Published: (2025)
Beyond Static Bias: Adaptive Multi-Fidelity Bandits with Improving Proxies
by: Lu, Muyun, et al.
Published: (2026)
by: Lu, Muyun, et al.
Published: (2026)
Have You Poisoned My Data? Defending Neural Networks against Data Poisoning
by: De Gaspari, Fabio, et al.
Published: (2024)
by: De Gaspari, Fabio, et al.
Published: (2024)
On Robustness of Linear Classifiers to Targeted Data Poisoning
by: Gupta, Nakshatra, et al.
Published: (2025)
by: Gupta, Nakshatra, et al.
Published: (2025)
When Do "More Contexts" Help with Sarcasm Recognition?
by: Nimase, Ojas, et al.
Published: (2024)
by: Nimase, Ojas, et al.
Published: (2024)
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
by: Wei, Yukuan, et al.
Published: (2025)
by: Wei, Yukuan, et al.
Published: (2025)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
by: Shakerinava, Mehran, et al.
Published: (2025)
by: Shakerinava, Mehran, et al.
Published: (2025)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
by: Boone, Victor, et al.
Published: (2024)
by: Boone, Victor, et al.
Published: (2024)
A Linear Approach to Data Poisoning
by: Flynn, Donald, et al.
Published: (2025)
by: Flynn, Donald, et al.
Published: (2025)
Design-Based Bandits Under Network Interference: Trade-Off Between Regret and Statistical Inference
by: Wang, Zichen, et al.
Published: (2025)
by: Wang, Zichen, et al.
Published: (2025)
Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Two-Timescale Critic-Actor for Average Reward MDPs with Function Approximation
by: Panda, Prashansa, et al.
Published: (2024)
by: Panda, Prashansa, et al.
Published: (2024)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
by: Gadot, Uri, et al.
Published: (2023)
by: Gadot, Uri, et al.
Published: (2023)
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
Confundo: Learning to Generate Robust Poison for Practical RAG Systems
by: Hu, Haoyang, et al.
Published: (2026)
by: Hu, Haoyang, et al.
Published: (2026)
Sample Complexity Characterization for Linear Contextual MDPs
by: Deng, Junze, et al.
Published: (2024)
by: Deng, Junze, et al.
Published: (2024)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach
by: Ganesh, Swetha, et al.
Published: (2025)
by: Ganesh, Swetha, et al.
Published: (2025)
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
by: Souly, Alexandra, et al.
Published: (2025)
by: Souly, Alexandra, et al.
Published: (2025)
FuncPoison: Poisoning Function Library to Hijack Multi-agent Autonomous Driving Systems
by: Long, Yuzhen, et al.
Published: (2025)
by: Long, Yuzhen, et al.
Published: (2025)
Span-Based Optimal Sample Complexity for Average Reward MDPs
by: Zurek, Matthew, et al.
Published: (2023)
by: Zurek, Matthew, et al.
Published: (2023)
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
by: Liu, Yinuo, et al.
Published: (2025)
by: Liu, Yinuo, et al.
Published: (2025)
Associative Poisoning to Generative Machine Learning
by: Mohus, Mathias Lundteigen, et al.
Published: (2025)
by: Mohus, Mathias Lundteigen, et al.
Published: (2025)
Similar Items
-
Hard Work Does Not Always Pay Off: Poisoning Attacks on Neural Architecture Search
by: Coalson, Zachary, et al.
Published: (2024) -
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024) -
Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
by: Sasnauskas, Paulius, et al.
Published: (2025) -
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2025) -
Efficient Reinforcement Learning in Probabilistic Reward Machines
by: Lin, Xiaofeng, et al.
Published: (2024)