Model approximation in MDPs with unbounded per-step cost

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bozkurt, Berk, Mahajan, Aditya, Nayyar, Ashutosh, Ouyang, Yi
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913233862918144
author Bozkurt, Berk
Mahajan, Aditya
Nayyar, Ashutosh
Ouyang, Yi
author_facet Bozkurt, Berk
Mahajan, Aditya
Nayyar, Ashutosh
Ouyang, Yi
contents We consider the problem of designing a control policy for an infinite-horizon discounted cost Markov decision process $\mathcal{M}$ when we only have access to an approximate model $\hat{\mathcal{M}}$. How well does an optimal policy $\hatπ^{\star}$ of the approximate model perform when used in the original model $\mathcal{M}$? We answer this question by bounding a weighted norm of the difference between the value function of $\hatπ^\star $ when used in $\mathcal{M}$ and the optimal value function of $\mathcal{M}$. We then extend our results and obtain potentially tighter upper bounds by considering affine transformations of the per-step cost. We further provide upper bounds that explicitly depend on the weighted distance between cost functions and weighted distance between transition kernels of the original and approximate models. We present examples to illustrate our results.
format Preprint
id arxiv_https___arxiv_org_abs_2402_08813
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Model approximation in MDPs with unbounded per-step cost
Bozkurt, Berk
Mahajan, Aditya
Nayyar, Ashutosh
Ouyang, Yi
Optimization and Control
Machine Learning
Systems and Control
We consider the problem of designing a control policy for an infinite-horizon discounted cost Markov decision process $\mathcal{M}$ when we only have access to an approximate model $\hat{\mathcal{M}}$. How well does an optimal policy $\hatπ^{\star}$ of the approximate model perform when used in the original model $\mathcal{M}$? We answer this question by bounding a weighted norm of the difference between the value function of $\hatπ^\star $ when used in $\mathcal{M}$ and the optimal value function of $\mathcal{M}$. We then extend our results and obtain potentially tighter upper bounds by considering affine transformations of the per-step cost. We further provide upper bounds that explicitly depend on the weighted distance between cost functions and weighted distance between transition kernels of the original and approximate models. We present examples to illustrate our results.
title Model approximation in MDPs with unbounded per-step cost
topic Optimization and Control
Machine Learning
Systems and Control
url https://arxiv.org/abs/2402.08813