Primal-Only Actor Critic Algorithm for Robust Constrained Average Cost MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917069245644800 |
|---|---|
| author | Satheesh, Anirudh Sathish, Sooraj Ganesh, Swetha Powell, Keenan Aggarwal, Vaneet |
| author_facet | Satheesh, Anirudh Sathish, Sooraj Ganesh, Swetha Powell, Keenan Aggarwal, Vaneet |
| contents | In this work, we study the problem of finding robust and safe policies in Robust Constrained Average-Cost Markov Decision Processes (RCMDPs). A key challenge in this setting is the lack of strong duality, which prevents the direct use of standard primal-dual methods for constrained RL. Additional difficulties arise from the average-cost setting, where the Robust Bellman operator is not a contraction under any norm. To address these challenges, we propose an actor-critic algorithm for Average-Cost RCMDPs. We show that our method achieves both \(ε\)-feasibility and \(ε\)-optimality, and we establish a sample complexities of \(\tilde{O}\left(ε^{-4}\right)\) and \(\tilde{O}\left(ε^{-6}\right)\) with and without slackness assumption, which is comparable to the discounted setting. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_05758 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Primal-Only Actor Critic Algorithm for Robust Constrained Average Cost MDPs Satheesh, Anirudh Sathish, Sooraj Ganesh, Swetha Powell, Keenan Aggarwal, Vaneet Machine Learning In this work, we study the problem of finding robust and safe policies in Robust Constrained Average-Cost Markov Decision Processes (RCMDPs). A key challenge in this setting is the lack of strong duality, which prevents the direct use of standard primal-dual methods for constrained RL. Additional difficulties arise from the average-cost setting, where the Robust Bellman operator is not a contraction under any norm. To address these challenges, we propose an actor-critic algorithm for Average-Cost RCMDPs. We show that our method achieves both \(ε\)-feasibility and \(ε\)-optimality, and we establish a sample complexities of \(\tilde{O}\left(ε^{-4}\right)\) and \(\tilde{O}\left(ε^{-6}\right)\) with and without slackness assumption, which is comparable to the discounted setting. |
| title | Primal-Only Actor Critic Algorithm for Robust Constrained Average Cost MDPs |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2511.05758 |