Planning in entropy-regularized Markov decision processes and games

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Grill, Jean-Bastien, Domingues, Omar Darwiche, Ménard, Pierre, Munos, Rémi, Valko, Michal
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914496356810752
author Grill, Jean-Bastien
Domingues, Omar Darwiche
Ménard, Pierre
Munos, Rémi
Valko, Michal
author_facet Grill, Jean-Bastien
Domingues, Omar Darwiche
Ménard, Pierre
Munos, Rémi
Valko, Michal
contents We propose SmoothCruiser, a new planning algorithm for estimating the value function in entropy-regularized Markov decision processes and two-player games, given a generative model of the environment. SmoothCruiser makes use of the smoothness of the Bellman operator promoted by the regularization to achieve problem-independent sample complexity of order O~(1/epsilon^4) for a desired accuracy epsilon, whereas for non-regularized settings there are no known algorithms with guaranteed polynomial sample complexity in the worst case.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19695
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Planning in entropy-regularized Markov decision processes and games
Grill, Jean-Bastien
Domingues, Omar Darwiche
Ménard, Pierre
Munos, Rémi
Valko, Michal
Machine Learning
We propose SmoothCruiser, a new planning algorithm for estimating the value function in entropy-regularized Markov decision processes and two-player games, given a generative model of the environment. SmoothCruiser makes use of the smoothness of the Bellman operator promoted by the regularization to achieve problem-independent sample complexity of order O~(1/epsilon^4) for a desired accuracy epsilon, whereas for non-regularized settings there are no known algorithms with guaranteed polynomial sample complexity in the worst case.
title Planning in entropy-regularized Markov decision processes and games
topic Machine Learning
url https://arxiv.org/abs/2604.19695