Off-Policy Reinforcement Learning with High Dimensional Reward

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Dong Neuck, Kosorok, Michael R.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916357394661376
author Lee, Dong Neuck
Kosorok, Michael R.
author_facet Lee, Dong Neuck
Kosorok, Michael R.
contents Conventional off-policy reinforcement learning (RL) focuses on maximizing the expected return of scalar rewards. Distributional RL (DRL), in contrast, studies the distribution of returns with the distributional Bellman operator in a Euclidean space, leading to highly flexible choices for utility. This paper establishes robust theoretical foundations for DRL. We prove the contraction property of the Bellman operator even when the reward space is an infinite-dimensional separable Banach space. Furthermore, we demonstrate that the behavior of high- or infinite-dimensional returns can be effectively approximated using a lower-dimensional Euclidean space. Leveraging these theoretical insights, we propose a novel DRL algorithm that tackles problems which have been previously intractable using conventional reinforcement learning approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2408_07660
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Off-Policy Reinforcement Learning with High Dimensional Reward
Lee, Dong Neuck
Kosorok, Michael R.
Machine Learning
68T05, 46B09 (Primary) 46B06 (Secondary)
Conventional off-policy reinforcement learning (RL) focuses on maximizing the expected return of scalar rewards. Distributional RL (DRL), in contrast, studies the distribution of returns with the distributional Bellman operator in a Euclidean space, leading to highly flexible choices for utility. This paper establishes robust theoretical foundations for DRL. We prove the contraction property of the Bellman operator even when the reward space is an infinite-dimensional separable Banach space. Furthermore, we demonstrate that the behavior of high- or infinite-dimensional returns can be effectively approximated using a lower-dimensional Euclidean space. Leveraging these theoretical insights, we propose a novel DRL algorithm that tackles problems which have been previously intractable using conventional reinforcement learning approaches.
title Off-Policy Reinforcement Learning with High Dimensional Reward
topic Machine Learning
68T05, 46B09 (Primary) 46B06 (Secondary)
url https://arxiv.org/abs/2408.07660