Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
Fuente:
arXiv
Saved in:
| Main Authors: | Ganguly, Sourav, Pandit, Kartik, Ghosh, Arnob |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Certifiable Safe RLHF: Fixed-Penalty Constraint Optimization for Safer Language Models
by: Pandit, Kartik, et al.
Published: (2025)
by: Pandit, Kartik, et al.
Published: (2025)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees
by: Zhang, Zuyuan, et al.
Published: (2025)
by: Zhang, Zuyuan, et al.
Published: (2025)
Optimistic Regret Bounds for Online Learning in Adversarial Markov Decision Processes
by: Moon, Sang Bin, et al.
Published: (2024)
by: Moon, Sang Bin, et al.
Published: (2024)
Provably Efficient Sample Complexity for Robust CMDP
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
Pessimistic Off-Policy Optimization for Learning to Rank
by: Cief, Matej, et al.
Published: (2022)
by: Cief, Matej, et al.
Published: (2022)
Optimistic Policy Regularization
by: Pham, Mai, et al.
Published: (2026)
by: Pham, Mai, et al.
Published: (2026)
Kernel-Based Function Approximation for Average Reward Reinforcement Learning: An Optimist No-Regret Algorithm
by: Vakili, Sattar, et al.
Published: (2024)
by: Vakili, Sattar, et al.
Published: (2024)
Optimistic Thompson Sampling for No-Regret Learning in Unknown Games
by: Li, Yingru, et al.
Published: (2024)
by: Li, Yingru, et al.
Published: (2024)
An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints
by: Zhu, Jiahui, et al.
Published: (2025)
by: Zhu, Jiahui, et al.
Published: (2025)
Regret-Based Defense in Adversarial Reinforcement Learning
by: Belaire, Roman, et al.
Published: (2023)
by: Belaire, Roman, et al.
Published: (2023)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
by: Zhai, Yuanzhao, et al.
Published: (2024)
by: Zhai, Yuanzhao, et al.
Published: (2024)
Data-Driven Online Model Selection With Regret Guarantees
by: Pacchiano, Aldo, et al.
Published: (2023)
by: Pacchiano, Aldo, et al.
Published: (2023)
Minimizing Weighted Counterfactual Regret with Optimistic Online Mirror Descent
by: Xu, Hang, et al.
Published: (2024)
by: Xu, Hang, et al.
Published: (2024)
Pessimistic Causal Reinforcement Learning with Mediators for Confounded Offline Data
by: Wang, Danyang, et al.
Published: (2024)
by: Wang, Danyang, et al.
Published: (2024)
Optimistic Reinforcement Learning with Quantile Objectives
by: Alipour-Vaezi, Mohammad, et al.
Published: (2025)
by: Alipour-Vaezi, Mohammad, et al.
Published: (2025)
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
by: Fluri, Lukas, et al.
Published: (2024)
by: Fluri, Lukas, et al.
Published: (2024)
Conservative DDPG -- Pessimistic RL without Ensemble
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Pessimistic Iterative Planning with RNNs for Robust POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2024)
by: Galesloot, Maris F. L., et al.
Published: (2024)
Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation
by: Jiang, Jie, et al.
Published: (2026)
by: Jiang, Jie, et al.
Published: (2026)
Optimistic Rates for Learning from Label Proportions
by: Li, Gene, et al.
Published: (2024)
by: Li, Gene, et al.
Published: (2024)
Pessimistic Value Iteration for Multi-Task Data Sharing in Offline Reinforcement Learning
by: Bai, Chenjia, et al.
Published: (2024)
by: Bai, Chenjia, et al.
Published: (2024)
Restricted Bayesian Neural Network
by: Ganguly, Sourav, et al.
Published: (2024)
by: Ganguly, Sourav, et al.
Published: (2024)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
by: McCarthy, James, et al.
Published: (2025)
by: McCarthy, James, et al.
Published: (2025)
Diverse Randomized Value Functions: A Provably Pessimistic Approach for Offline Reinforcement Learning
by: Yu, Xudong, et al.
Published: (2024)
by: Yu, Xudong, et al.
Published: (2024)
Learning in Markov Games with Adaptive Adversaries: Policy Regret, Fundamental Barriers, and Efficient Algorithms
by: Nguyen-Tang, Thanh, et al.
Published: (2024)
by: Nguyen-Tang, Thanh, et al.
Published: (2024)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
by: Wu, Yongtao, et al.
Published: (2025)
by: Wu, Yongtao, et al.
Published: (2025)
Learning Optimal and Sample-Efficient Decision Policies with Guarantees
by: Shao, Daqian
Published: (2026)
by: Shao, Daqian
Published: (2026)
Adversarial Environment Design via Regret-Guided Diffusion Models
by: Chung, Hojun, et al.
Published: (2024)
by: Chung, Hojun, et al.
Published: (2024)
Quantum Speedups in Regret Analysis of Infinite Horizon Average-Reward Markov Decision Processes
by: Ganguly, Bhargav, et al.
Published: (2023)
by: Ganguly, Bhargav, et al.
Published: (2023)
No-Regret Learning Under Adversarial Resource Constraints: A Spending Plan Is All You Need!
by: Stradi, Francesco Emanuele, et al.
Published: (2025)
by: Stradi, Francesco Emanuele, et al.
Published: (2025)
Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity
by: Arnob, Samin Yeasar, et al.
Published: (2025)
by: Arnob, Samin Yeasar, et al.
Published: (2025)
Optimistic Gradient Learning with Hessian Corrections for High-Dimensional Black-Box Optimization
by: Kfir, Yedidya, et al.
Published: (2025)
by: Kfir, Yedidya, et al.
Published: (2025)
Statistical Guarantees in Synthetic Data through Conformal Adversarial Generation
by: Vishwakarma, Rahul, et al.
Published: (2025)
by: Vishwakarma, Rahul, et al.
Published: (2025)
Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data
by: Xue, Ruiqi, et al.
Published: (2026)
by: Xue, Ruiqi, et al.
Published: (2026)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Implementing Quantum Generative Adversarial Network (qGAN) and QCBM in Finance
by: Ganguly, Santanu
Published: (2023)
by: Ganguly, Santanu
Published: (2023)
Guide: Generalized-Prior and Data Encoders for DAG Estimation
by: Roy, Amartya, et al.
Published: (2025)
by: Roy, Amartya, et al.
Published: (2025)
Deep Generative Models for Offline Policy Learning: Tutorial, Survey, and Perspectives on Future Directions
by: Chen, Jiayu, et al.
Published: (2024)
by: Chen, Jiayu, et al.
Published: (2024)
Similar Items
-
Certifiable Safe RLHF: Fixed-Penalty Constraint Optimization for Safer Language Models
by: Pandit, Kartik, et al.
Published: (2025) -
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025) -
Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees
by: Zhang, Zuyuan, et al.
Published: (2025) -
Optimistic Regret Bounds for Online Learning in Adversarial Markov Decision Processes
by: Moon, Sang Bin, et al.
Published: (2024) -
Provably Efficient Sample Complexity for Robust CMDP
by: Ganguly, Sourav, et al.
Published: (2025)