Maximum entropy GFlowNets with soft Q-learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Mohammadpour, Sobhan, Bengio, Emmanuel, Frejinger, Emma, Bacon, Pierre-Luc
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910838416211968
author Mohammadpour, Sobhan
Bengio, Emmanuel
Frejinger, Emma
Bacon, Pierre-Luc
author_facet Mohammadpour, Sobhan
Bengio, Emmanuel
Frejinger, Emma
Bacon, Pierre-Luc
contents Generative Flow Networks (GFNs) have emerged as a powerful tool for sampling discrete objects from unnormalized distributions, offering a scalable alternative to Markov Chain Monte Carlo (MCMC) methods. While GFNs draw inspiration from maximum entropy reinforcement learning (RL), the connection between the two has largely been unclear and seemingly applicable only in specific cases. This paper addresses the connection by constructing an appropriate reward function, thereby establishing an exact relationship between GFNs and maximum entropy RL. This construction allows us to introduce maximum entropy GFNs, which, in contrast to GFNs with uniform backward policy, achieve the maximum entropy attainable by GFNs without constraints on the state space.
format Preprint
id arxiv_https___arxiv_org_abs_2312_14331
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Maximum entropy GFlowNets with soft Q-learning
Mohammadpour, Sobhan
Bengio, Emmanuel
Frejinger, Emma
Bacon, Pierre-Luc
Machine Learning
Generative Flow Networks (GFNs) have emerged as a powerful tool for sampling discrete objects from unnormalized distributions, offering a scalable alternative to Markov Chain Monte Carlo (MCMC) methods. While GFNs draw inspiration from maximum entropy reinforcement learning (RL), the connection between the two has largely been unclear and seemingly applicable only in specific cases. This paper addresses the connection by constructing an appropriate reward function, thereby establishing an exact relationship between GFNs and maximum entropy RL. This construction allows us to introduce maximum entropy GFNs, which, in contrast to GFNs with uniform backward policy, achieve the maximum entropy attainable by GFNs without constraints on the state space.
title Maximum entropy GFlowNets with soft Q-learning
topic Machine Learning
url https://arxiv.org/abs/2312.14331