Rethinking Adversarial Policies: A Generalized Attack Formulation and Provable Defense in RL
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xiangyu, Chakraborty, Souradip, Sun, Yanchao, Huang, Furong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Worst-case Attacks: Robust RL with Adaptive Defense via Non-dominated Policies
von: Liu, Xiangyu, et al.
Veröffentlicht: (2024)
von: Liu, Xiangyu, et al.
Veröffentlicht: (2024)
Rethinking Adversarial Attacks in Reinforcement Learning from Policy Distribution Perspective
von: Duan, Tianyang, et al.
Veröffentlicht: (2025)
von: Duan, Tianyang, et al.
Veröffentlicht: (2025)
Agentic Critical Training
von: Liu, Weize, et al.
Veröffentlicht: (2026)
von: Liu, Weize, et al.
Veröffentlicht: (2026)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
von: Barakat, Anas, et al.
Veröffentlicht: (2024)
von: Barakat, Anas, et al.
Veröffentlicht: (2024)
Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled Perturbations
von: Liang, Yongyuan, et al.
Veröffentlicht: (2023)
von: Liang, Yongyuan, et al.
Veröffentlicht: (2023)
Towards Robust Policy: Enhancing Offline Reinforcement Learning with Adversarial Attacks and Defenses
von: Nguyen, Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thanh, et al.
Veröffentlicht: (2024)
Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation
von: Shang, Yingjia, et al.
Veröffentlicht: (2025)
von: Shang, Yingjia, et al.
Veröffentlicht: (2025)
Adversarial Training for Defense Against Label Poisoning Attacks
von: Bal, Melis Ilayda, et al.
Veröffentlicht: (2025)
von: Bal, Melis Ilayda, et al.
Veröffentlicht: (2025)
Model Mimic Attack: Knowledge Distillation for Provably Transferable Adversarial Examples
von: Lukyanov, Kirill, et al.
Veröffentlicht: (2024)
von: Lukyanov, Kirill, et al.
Veröffentlicht: (2024)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
Rethinking the Intermediate Features in Adversarial Attacks: Misleading Robotic Models via Adversarial Distillation
von: Zhao, Ke, et al.
Veröffentlicht: (2024)
von: Zhao, Ke, et al.
Veröffentlicht: (2024)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
Provably Invincible Adversarial Attacks on Reinforcement Learning Systems: A Rate-Distortion Information-Theoretic Approach
von: Lu, Ziqing, et al.
Veröffentlicht: (2025)
von: Lu, Ziqing, et al.
Veröffentlicht: (2025)
Deep Adversarial Defense Against Multilevel-Lp Attacks
von: Wang, Ren, et al.
Veröffentlicht: (2024)
von: Wang, Ren, et al.
Veröffentlicht: (2024)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
von: Barakat, Anas, et al.
Veröffentlicht: (2026)
von: Barakat, Anas, et al.
Veröffentlicht: (2026)
MaxMin-RLHF: Alignment with Diverse Human Preferences
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
Enhancing Security in Deep Reinforcement Learning: A Comprehensive Survey on Adversarial Attacks and Defenses
von: Yichao, Wu, et al.
Veröffentlicht: (2025)
von: Yichao, Wu, et al.
Veröffentlicht: (2025)
On Minimizing Adversarial Counterfactual Error in Adversarial RL
von: Belaire, Roman, et al.
Veröffentlicht: (2024)
von: Belaire, Roman, et al.
Veröffentlicht: (2024)
TACO: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement Learning
von: Zheng, Ruijie, et al.
Veröffentlicht: (2023)
von: Zheng, Ruijie, et al.
Veröffentlicht: (2023)
Deep RL With Information Constrained Policies: Generalization in Continuous Control
von: Malloy, Tailia, et al.
Veröffentlicht: (2020)
von: Malloy, Tailia, et al.
Veröffentlicht: (2020)
Time-Constrained Recommendations: Reinforcement Learning Strategies for E-Commerce
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2025)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2025)
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
von: Liu, Zhihan, et al.
Veröffentlicht: (2024)
von: Liu, Zhihan, et al.
Veröffentlicht: (2024)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
von: Kim, Juno, et al.
Veröffentlicht: (2025)
von: Kim, Juno, et al.
Veröffentlicht: (2025)
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
von: Ran-Milo, Yuval, et al.
Veröffentlicht: (2026)
von: Ran-Milo, Yuval, et al.
Veröffentlicht: (2026)
DISC: Decoupling Instruction from State-Conditioned Control via Policy Generation
von: Ren, Hanxiang, et al.
Veröffentlicht: (2026)
von: Ren, Hanxiang, et al.
Veröffentlicht: (2026)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
Fast Adversarial Training against Sparse Attacks Requires Loss Smoothing
von: Zhong, Xuyang, et al.
Veröffentlicht: (2025)
von: Zhong, Xuyang, et al.
Veröffentlicht: (2025)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
Defense without Forgetting: Continual Adversarial Defense with Anisotropic & Isotropic Pseudo Replay
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
Towards Provable Log Density Policy Gradient
von: Katdare, Pulkit, et al.
Veröffentlicht: (2024)
von: Katdare, Pulkit, et al.
Veröffentlicht: (2024)
Is poisoning a real threat to LLM alignment? Maybe more so than you think
von: Pathmanathan, Pankayaraj, et al.
Veröffentlicht: (2024)
von: Pathmanathan, Pankayaraj, et al.
Veröffentlicht: (2024)
TabAttackBench: A Benchmark for Adversarial Attacks on Tabular Data
von: He, Zhipeng, et al.
Veröffentlicht: (2025)
von: He, Zhipeng, et al.
Veröffentlicht: (2025)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model
von: Pathmanathan, Pankayaraj, et al.
Veröffentlicht: (2026)
von: Pathmanathan, Pankayaraj, et al.
Veröffentlicht: (2026)
Regret-Based Defense in Adversarial Reinforcement Learning
von: Belaire, Roman, et al.
Veröffentlicht: (2023)
von: Belaire, Roman, et al.
Veröffentlicht: (2023)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2024)
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2024)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
von: Trivedi, Prashant, et al.
Veröffentlicht: (2025)
von: Trivedi, Prashant, et al.
Veröffentlicht: (2025)
OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
von: Luo, Yu, et al.
Veröffentlicht: (2024)
von: Luo, Yu, et al.
Veröffentlicht: (2024)
Attacks and Defenses for Generative Diffusion Models: A Comprehensive Survey
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2024)
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Worst-case Attacks: Robust RL with Adaptive Defense via Non-dominated Policies
von: Liu, Xiangyu, et al.
Veröffentlicht: (2024) -
Rethinking Adversarial Attacks in Reinforcement Learning from Policy Distribution Perspective
von: Duan, Tianyang, et al.
Veröffentlicht: (2025) -
Agentic Critical Training
von: Liu, Weize, et al.
Veröffentlicht: (2026) -
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
von: Barakat, Anas, et al.
Veröffentlicht: (2024) -
Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled Perturbations
von: Liang, Yongyuan, et al.
Veröffentlicht: (2023)