Decision Flow Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Jifeng, Huang, Sili, Guo, Siyuan, Liu, Zhaogeng, Shen, Li, Sun, Lichao, Chen, Hechang, Chang, Yi, Tao, Dacheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916760227151872
author Hu, Jifeng
Huang, Sili
Guo, Siyuan
Liu, Zhaogeng
Shen, Li
Sun, Lichao
Chen, Hechang
Chang, Yi
Tao, Dacheng
author_facet Hu, Jifeng
Huang, Sili
Guo, Siyuan
Liu, Zhaogeng
Shen, Li
Sun, Lichao
Chen, Hechang
Chang, Yi
Tao, Dacheng
contents In recent years, generative models have shown remarkable capabilities across diverse fields, including images, videos, language, and decision-making. By applying powerful generative models such as flow-based models to reinforcement learning, we can effectively model complex multi-modal action distributions and achieve superior robotic control in continuous action spaces, surpassing the limitations of single-modal action distributions with traditional Gaussian-based policies. Previous methods usually adopt the generative models as behavior models to fit state-conditioned action distributions from datasets, with policy optimization conducted separately through additional policies using value-based sample weighting or gradient-based updates. However, this separation prevents the simultaneous optimization of multi-modal distribution fitting and policy improvement, ultimately hindering the training of models and degrading the performance. To address this issue, we propose Decision Flow, a unified framework that integrates multi-modal action distribution modeling and policy optimization. Specifically, our method formulates the action generation procedure of flow-based models as a flow decision-making process, where each action generation step corresponds to one flow decision. Consequently, our method seamlessly optimizes the flow policy while capturing multi-modal action distributions. We provide rigorous proofs of Decision Flow and validate the effectiveness through extensive experiments across dozens of offline RL environments. Compared with established offline RL baselines, the results demonstrate that our method achieves or matches the SOTA performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20350
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Decision Flow Policy Optimization
Hu, Jifeng
Huang, Sili
Guo, Siyuan
Liu, Zhaogeng
Shen, Li
Sun, Lichao
Chen, Hechang
Chang, Yi
Tao, Dacheng
Machine Learning
Artificial Intelligence
In recent years, generative models have shown remarkable capabilities across diverse fields, including images, videos, language, and decision-making. By applying powerful generative models such as flow-based models to reinforcement learning, we can effectively model complex multi-modal action distributions and achieve superior robotic control in continuous action spaces, surpassing the limitations of single-modal action distributions with traditional Gaussian-based policies. Previous methods usually adopt the generative models as behavior models to fit state-conditioned action distributions from datasets, with policy optimization conducted separately through additional policies using value-based sample weighting or gradient-based updates. However, this separation prevents the simultaneous optimization of multi-modal distribution fitting and policy improvement, ultimately hindering the training of models and degrading the performance. To address this issue, we propose Decision Flow, a unified framework that integrates multi-modal action distribution modeling and policy optimization. Specifically, our method formulates the action generation procedure of flow-based models as a flow decision-making process, where each action generation step corresponds to one flow decision. Consequently, our method seamlessly optimizes the flow policy while capturing multi-modal action distributions. We provide rigorous proofs of Decision Flow and validate the effectiveness through extensive experiments across dozens of offline RL environments. Compared with established offline RL baselines, the results demonstrate that our method achieves or matches the SOTA performance.
title Decision Flow Policy Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.20350