Parameterized Projected Bellman Operator

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Vincent, Théo, Metelli, Alberto Maria, Belousov, Boris, Peters, Jan, Restelli, Marcello, D'Eramo, Carlo
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916148358938624
author Vincent, Théo
Metelli, Alberto Maria
Belousov, Boris
Peters, Jan
Restelli, Marcello
D'Eramo, Carlo
author_facet Vincent, Théo
Metelli, Alberto Maria
Belousov, Boris
Peters, Jan
Restelli, Marcello
D'Eramo, Carlo
contents Approximate value iteration (AVI) is a family of algorithms for reinforcement learning (RL) that aims to obtain an approximation of the optimal value function. Generally, AVI algorithms implement an iterated procedure where each step consists of (i) an application of the Bellman operator and (ii) a projection step into a considered function space. Notoriously, the Bellman operator leverages transition samples, which strongly determine its behavior, as uninformative samples can result in negligible updates or long detours, whose detrimental effects are further exacerbated by the computationally intensive projection step. To address these issues, we propose a novel alternative approach based on learning an approximate version of the Bellman operator rather than estimating it through samples as in AVI approaches. This way, we are able to (i) generalize across transition samples and (ii) avoid the computationally intensive projection step. For this reason, we call our novel operator projected Bellman operator (PBO). We formulate an optimization problem to learn PBO for generic sequential decision-making problems, and we theoretically analyze its properties in two representative classes of RL problems. Furthermore, we theoretically study our approach under the lens of AVI and devise algorithmic implementations to learn PBO in offline and online settings by leveraging neural network parameterizations. Finally, we empirically showcase the benefits of PBO w.r.t. the regular Bellman operator on several RL problems.
format Preprint
id arxiv_https___arxiv_org_abs_2312_12869
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Parameterized Projected Bellman Operator
Vincent, Théo
Metelli, Alberto Maria
Belousov, Boris
Peters, Jan
Restelli, Marcello
D'Eramo, Carlo
Machine Learning
Artificial Intelligence
Approximate value iteration (AVI) is a family of algorithms for reinforcement learning (RL) that aims to obtain an approximation of the optimal value function. Generally, AVI algorithms implement an iterated procedure where each step consists of (i) an application of the Bellman operator and (ii) a projection step into a considered function space. Notoriously, the Bellman operator leverages transition samples, which strongly determine its behavior, as uninformative samples can result in negligible updates or long detours, whose detrimental effects are further exacerbated by the computationally intensive projection step. To address these issues, we propose a novel alternative approach based on learning an approximate version of the Bellman operator rather than estimating it through samples as in AVI approaches. This way, we are able to (i) generalize across transition samples and (ii) avoid the computationally intensive projection step. For this reason, we call our novel operator projected Bellman operator (PBO). We formulate an optimization problem to learn PBO for generic sequential decision-making problems, and we theoretically analyze its properties in two representative classes of RL problems. Furthermore, we theoretically study our approach under the lens of AVI and devise algorithmic implementations to learn PBO in offline and online settings by leveraging neural network parameterizations. Finally, we empirically showcase the benefits of PBO w.r.t. the regular Bellman operator on several RL problems.
title Parameterized Projected Bellman Operator
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2312.12869