Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
Fuente:
arXiv
Saved in:
| Main Authors: | Ganguly, Sourav, Panaganti, Kishan, Ghosh, Arnob, Wierman, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Provably Efficient Sample Complexity for Robust CMDP
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
Certifiable Safe RLHF: Fixed-Penalty Constraint Optimization for Safer Language Models
by: Pandit, Kartik, et al.
Published: (2025)
by: Pandit, Kartik, et al.
Published: (2025)
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
by: Ganguly, Sourav, et al.
Published: (2026)
by: Ganguly, Sourav, et al.
Published: (2026)
Multi-CALF: A Policy Combination Approach with Statistical Guarantees
by: Malaniya, Georgiy, et al.
Published: (2025)
by: Malaniya, Georgiy, et al.
Published: (2025)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
Safety Optimized Reinforcement Learning via Multi-Objective Policy Optimization
by: Honari, Homayoun, et al.
Published: (2024)
by: Honari, Homayoun, et al.
Published: (2024)
Scaling Learning based Policy Optimization for Temporal Logic Tasks by Controller Network Dropout
by: Hashemi, Navid, et al.
Published: (2024)
by: Hashemi, Navid, et al.
Published: (2024)
Adaptive Smooth Tchebycheff Attention for Multi-Objective Policy Optimization
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2026)
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2026)
Frugal Actor-Critic: Sample Efficient Off-Policy Deep Reinforcement Learning Using Unique Experiences
by: Singh, Nikhil Kumar, et al.
Published: (2024)
by: Singh, Nikhil Kumar, et al.
Published: (2024)
Probabilistic Constrained Reinforcement Learning with Formal Interpretability
by: Wang, Yanran, et al.
Published: (2023)
by: Wang, Yanran, et al.
Published: (2023)
Vegetable Peeling: A Case Study in Constrained Dexterous Manipulation
by: Chen, Tao, et al.
Published: (2024)
by: Chen, Tao, et al.
Published: (2024)
Scalable Data-Driven Reachability Analysis and Control via Koopman Operators with Conformal Coverage Guarantees
by: Nath, Devesh, et al.
Published: (2026)
by: Nath, Devesh, et al.
Published: (2026)
Constrained Control for Autonomous Spacecraft Rendezvous: Learning-Based Time Shift Governor
by: Kim, Taehyeun, et al.
Published: (2024)
by: Kim, Taehyeun, et al.
Published: (2024)
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
by: Omidi, Saber, et al.
Published: (2025)
by: Omidi, Saber, et al.
Published: (2025)
Embedding Morphology into Transformers for Cross-Robot Policy Learning
by: Suzuki, Kei, et al.
Published: (2026)
by: Suzuki, Kei, et al.
Published: (2026)
Align and Filter: Improving Performance in Asynchronous On-Policy RL
by: Honari, Homayoun, et al.
Published: (2026)
by: Honari, Homayoun, et al.
Published: (2026)
Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving
by: Zhang, Zhihao, et al.
Published: (2025)
by: Zhang, Zhihao, et al.
Published: (2025)
Predictive Red Teaming: Breaking Policies Without Breaking Robots
by: Majumdar, Anirudha, et al.
Published: (2025)
by: Majumdar, Anirudha, et al.
Published: (2025)
Distributionally Robust Cooperative Multi-Agent Reinforcement Learning via Robust Value Factorization
by: Qu, Chengrui, et al.
Published: (2026)
by: Qu, Chengrui, et al.
Published: (2026)
Explainable Representation of Finite-Memory Policies for POMDPs using Decision Trees
by: Azeem, Muqsit, et al.
Published: (2024)
by: Azeem, Muqsit, et al.
Published: (2024)
Optimal Control-Based Baseline for Guided Exploration in Policy Gradient Methods
by: Lyu, Xubo, et al.
Published: (2020)
by: Lyu, Xubo, et al.
Published: (2020)
Learning Causal Structure Distributions for Robust Planning
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2025)
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2025)
KL-regularization Itself is Differentially Private in Bandits and RLHF
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
ManyQuadrupeds: Learning a Single Locomotion Policy for Diverse Quadruped Robots
by: Shafiee, Milad, et al.
Published: (2023)
by: Shafiee, Milad, et al.
Published: (2023)
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
by: Vasan, Gautham, et al.
Published: (2024)
by: Vasan, Gautham, et al.
Published: (2024)
Deep Model Predictive Optimization
by: Sacks, Jacob, et al.
Published: (2023)
by: Sacks, Jacob, et al.
Published: (2023)
Formal Synthesis of Certifiably Robust Neural Lyapunov-Barrier Certificates
by: Wang, Chengxiao, et al.
Published: (2026)
by: Wang, Chengxiao, et al.
Published: (2026)
GUIDEd Agents: Enhancing Navigation Policies through Task-Specific Uncertainty Abstraction in Localization-Limited Environments
by: Puthumanaillam, Gokul, et al.
Published: (2024)
by: Puthumanaillam, Gokul, et al.
Published: (2024)
Robust Recovery Controller for a Quadrupedal Robot using Deep Reinforcement Learning
by: Lee, Joonho, et al.
Published: (2019)
by: Lee, Joonho, et al.
Published: (2019)
Learning Multiple Initial Solutions to Optimization Problems
by: Sharony, Elad, et al.
Published: (2024)
by: Sharony, Elad, et al.
Published: (2024)
Architecture Is All You Need: Diversity-Enabled Sweet Spots for Robust Humanoid Locomotion
by: Werner, Blake, et al.
Published: (2025)
by: Werner, Blake, et al.
Published: (2025)
One Filter to Deploy Them All: Robust Safety for Quadrupedal Navigation in Unknown Environments
by: Lin, Albert, et al.
Published: (2024)
by: Lin, Albert, et al.
Published: (2024)
Constraint-Generation Policy Optimization (CGPO): Nonlinear Programming for Policy Optimization in Mixed Discrete-Continuous MDPs
by: Gimelfarb, Michael, et al.
Published: (2024)
by: Gimelfarb, Michael, et al.
Published: (2024)
An Integrated Imitation and Reinforcement Learning Methodology for Robust Agile Aircraft Control with Limited Pilot Demonstration Data
by: Sever, Gulay Goktas, et al.
Published: (2023)
by: Sever, Gulay Goktas, et al.
Published: (2023)
Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing
by: Zongo, Alex, et al.
Published: (2026)
by: Zongo, Alex, et al.
Published: (2026)
Temporal Transfer Learning for Traffic Optimization with Coarse-grained Advisory Autonomy
by: Cho, Jung-Hoon, et al.
Published: (2023)
by: Cho, Jung-Hoon, et al.
Published: (2023)
Verification of Neural Reachable Tubes via Scenario Optimization and Conformal Prediction
by: Lin, Albert, et al.
Published: (2023)
by: Lin, Albert, et al.
Published: (2023)
Control-ITRA: Controlling the Behavior of a Driving Model
by: Lioutas, Vasileios, et al.
Published: (2025)
by: Lioutas, Vasileios, et al.
Published: (2025)
From Imitation to Optimization: A Comparative Study of Offline Learning for Autonomous Driving
by: Guillen-Perez, Antonio
Published: (2025)
by: Guillen-Perez, Antonio
Published: (2025)
Learning Pivoting Manipulation with Force and Vision Feedback Using Optimization-based Demonstrations
by: Shirai, Yuki, et al.
Published: (2025)
by: Shirai, Yuki, et al.
Published: (2025)
Similar Items
-
Provably Efficient Sample Complexity for Robust CMDP
by: Ganguly, Sourav, et al.
Published: (2025) -
Certifiable Safe RLHF: Fixed-Penalty Constraint Optimization for Safer Language Models
by: Pandit, Kartik, et al.
Published: (2025) -
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
by: Ganguly, Sourav, et al.
Published: (2026) -
Multi-CALF: A Policy Combination Approach with Statistical Guarantees
by: Malaniya, Georgiy, et al.
Published: (2025) -
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
by: Kwon, Jeongyeol, et al.
Published: (2024)