MDP Geometry, Normalization and Reward Balancing Solvers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mustafin, Arsenii, Pakharev, Aleksei, Olshevsky, Alex, Paschalidis, Ioannis Ch.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912259205234688
author Mustafin, Arsenii
Pakharev, Aleksei
Olshevsky, Alex
Paschalidis, Ioannis Ch.
author_facet Mustafin, Arsenii
Pakharev, Aleksei
Olshevsky, Alex
Paschalidis, Ioannis Ch.
contents We present a new geometric interpretation of Markov Decision Processes (MDPs) with a natural normalization procedure that allows us to adjust the value function at each state without altering the advantage of any action with respect to any policy. This advantage-preserving transformation of the MDP motivates a class of algorithms which we call Reward Balancing, which solve MDPs by iterating through these transformations, until an approximately optimal policy can be trivially found. We provide a convergence analysis of several algorithms in this class, in particular showing that for MDPs for unknown transition probabilities we can improve upon state-of-the-art sample complexity results.
format Preprint
id arxiv_https___arxiv_org_abs_2407_06712
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MDP Geometry, Normalization and Reward Balancing Solvers
Mustafin, Arsenii
Pakharev, Aleksei
Olshevsky, Alex
Paschalidis, Ioannis Ch.
Machine Learning
Optimization and Control
We present a new geometric interpretation of Markov Decision Processes (MDPs) with a natural normalization procedure that allows us to adjust the value function at each state without altering the advantage of any action with respect to any policy. This advantage-preserving transformation of the MDP motivates a class of algorithms which we call Reward Balancing, which solve MDPs by iterating through these transformations, until an approximately optimal policy can be trivially found. We provide a convergence analysis of several algorithms in this class, in particular showing that for MDPs for unknown transition probabilities we can improve upon state-of-the-art sample complexity results.
title MDP Geometry, Normalization and Reward Balancing Solvers
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2407.06712