Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Lin, Drapeau, Samuel, Shao, Fanghao, Zhu, Xuekai, Xue, Bo, Song, Yunchong, Laurière, Mathieu, Lin, Zhouhan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910033413931008
author Chen, Lin
Drapeau, Samuel
Shao, Fanghao
Zhu, Xuekai
Xue, Bo
Song, Yunchong
Laurière, Mathieu
Lin, Zhouhan
author_facet Chen, Lin
Drapeau, Samuel
Shao, Fanghao
Zhu, Xuekai
Xue, Bo
Song, Yunchong
Laurière, Mathieu
Lin, Zhouhan
contents Generative Flow Network (GFlowNet) objectives implicitly fix an equal mixing of forward and backward policies, potentially constraining the exploration-exploitation trade-off during training. By further exploring the link between GFlowNets and Markov chains, we establish an equivalence between GFlowNet objectives and Markov chain reversibility, thereby revealing the origin of such constraints, and provide a framework for adapting Markov chain properties to GFlowNets. Building on these theoretical findings, we propose $α$-GFNs, which generalize the mixing via a tunable parameter $α$. This generalization enables direct control over exploration-exploitation dynamics to enhance mode discovery capabilities, while ensuring convergence to unique flows. Across various benchmarks, including Set, Bit Sequence, and Molecule Generation, $α$-GFN objectives consistently outperform previous GFlowNet objectives, achieving up to a $10 \times$ increase in the number of discovered modes.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01749
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives
Chen, Lin
Drapeau, Samuel
Shao, Fanghao
Zhu, Xuekai
Xue, Bo
Song, Yunchong
Laurière, Mathieu
Lin, Zhouhan
Artificial Intelligence
Machine Learning
Generative Flow Network (GFlowNet) objectives implicitly fix an equal mixing of forward and backward policies, potentially constraining the exploration-exploitation trade-off during training. By further exploring the link between GFlowNets and Markov chains, we establish an equivalence between GFlowNet objectives and Markov chain reversibility, thereby revealing the origin of such constraints, and provide a framework for adapting Markov chain properties to GFlowNets. Building on these theoretical findings, we propose $α$-GFNs, which generalize the mixing via a tunable parameter $α$. This generalization enables direct control over exploration-exploitation dynamics to enhance mode discovery capabilities, while ensuring convergence to unique flows. Across various benchmarks, including Set, Bit Sequence, and Molecule Generation, $α$-GFN objectives consistently outperform previous GFlowNet objectives, achieving up to a $10 \times$ increase in the number of discovered modes.
title Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.01749