Epsilon-Optimal Policies for Average-Cost Separable MDPs with Perturbations
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866918173153951744 |
|---|---|
| author | Kantawala, Dhairya |
| author_facet | Kantawala, Dhairya |
| contents | We study a class of infinite-horizon average-cost Markov Decision Processes (MDPs) whose reward and transition structures are nearly separable. For the totally separable baseline (that is, with no perturbation), we derive an explicit stationary decision rule that is exactly average-optimal. We then show that under an epsilon-perturbation of the separable structure, this policy remains epsilon-optimal, meaning that the loss in the average reward is of order O(epsilon). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_23335 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Epsilon-Optimal Policies for Average-Cost Separable MDPs with Perturbations Kantawala, Dhairya Optimization and Control We study a class of infinite-horizon average-cost Markov Decision Processes (MDPs) whose reward and transition structures are nearly separable. For the totally separable baseline (that is, with no perturbation), we derive an explicit stationary decision rule that is exactly average-optimal. We then show that under an epsilon-perturbation of the separable structure, this policy remains epsilon-optimal, meaning that the loss in the average reward is of order O(epsilon). |
| title | Epsilon-Optimal Policies for Average-Cost Separable MDPs with Perturbations |
| topic | Optimization and Control |
| url | https://arxiv.org/abs/2510.23335 |