Flow-Based Policy for Online Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lv, Lei, Li, Yunfei, Luo, Yu, Sun, Fuchun, Kong, Tao, Xu, Jiafeng, Ma, Xiao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916794886782976
author Lv, Lei
Li, Yunfei
Luo, Yu
Sun, Fuchun
Kong, Tao
Xu, Jiafeng
Ma, Xiao
author_facet Lv, Lei
Li, Yunfei
Luo, Yu
Sun, Fuchun
Kong, Tao
Xu, Jiafeng
Ma, Xiao
contents We present \textbf{FlowRL}, a novel framework for online reinforcement learning that integrates flow-based policy representation with Wasserstein-2-regularized optimization. We argue that in addition to training signals, enhancing the expressiveness of the policy class is crucial for the performance gains in RL. Flow-based generative models offer such potential, excelling at capturing complex, multimodal action distributions. However, their direct application in online RL is challenging due to a fundamental objective mismatch: standard flow training optimizes for static data imitation, while RL requires value-based policy optimization through a dynamic buffer, leading to difficult optimization landscapes. FlowRL first models policies via a state-dependent velocity field, generating actions through deterministic ODE integration from noise. We derive a constrained policy search objective that jointly maximizes Q through the flow policy while bounding the Wasserstein-2 distance to a behavior-optimal policy implicitly derived from the replay buffer. This formulation effectively aligns the flow optimization with the RL objective, enabling efficient and value-aware policy learning despite the complexity of the policy class. Empirical evaluations on DMControl and Humanoidbench demonstrate that FlowRL achieves competitive performance in online reinforcement learning benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12811
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Flow-Based Policy for Online Reinforcement Learning
Lv, Lei
Li, Yunfei
Luo, Yu
Sun, Fuchun
Kong, Tao
Xu, Jiafeng
Ma, Xiao
Machine Learning
Artificial Intelligence
We present \textbf{FlowRL}, a novel framework for online reinforcement learning that integrates flow-based policy representation with Wasserstein-2-regularized optimization. We argue that in addition to training signals, enhancing the expressiveness of the policy class is crucial for the performance gains in RL. Flow-based generative models offer such potential, excelling at capturing complex, multimodal action distributions. However, their direct application in online RL is challenging due to a fundamental objective mismatch: standard flow training optimizes for static data imitation, while RL requires value-based policy optimization through a dynamic buffer, leading to difficult optimization landscapes. FlowRL first models policies via a state-dependent velocity field, generating actions through deterministic ODE integration from noise. We derive a constrained policy search objective that jointly maximizes Q through the flow policy while bounding the Wasserstein-2 distance to a behavior-optimal policy implicitly derived from the replay buffer. This formulation effectively aligns the flow optimization with the RL objective, enabling efficient and value-aware policy learning despite the complexity of the policy class. Empirical evaluations on DMControl and Humanoidbench demonstrate that FlowRL achieves competitive performance in online reinforcement learning benchmarks.
title Flow-Based Policy for Online Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.12811