Saved in:
Bibliographic Details
Main Authors: Hao, Ruijie, Zhang, Longfei, Dai, Yang, Ma, Yang, Liang, Xingxing, Cheng, Guangquan
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.00977
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912995710337024
author Hao, Ruijie
Zhang, Longfei
Dai, Yang
Ma, Yang
Liang, Xingxing
Cheng, Guangquan
author_facet Hao, Ruijie
Zhang, Longfei
Dai, Yang
Ma, Yang
Liang, Xingxing
Cheng, Guangquan
contents Reinforcement Learning (RL) has proven highly effective in addressing complex control and decision-making tasks. However, in most traditional RL algorithms, the policy is typically parameterized as a diagonal Gaussian distribution, which constrains the policy from capturing multimodal distributions, making it difficult to cover the full range of optimal solutions in multi-solution problems, and the return is reduced to a mean value, losing its multimodal nature and thus providing insufficient guidance for policy updates. In response to these problems, we propose a RL algorithm termed flow-based policy with distributional RL (FP-DRL). This algorithm models the policy using flow matching, which offers both computational efficiency and the capacity to fit complex distributions. Additionally, it employs a distributional RL approach to model and optimize the entire return distribution, thereby more effectively guiding multimodal policy updates and improving agent performance. Experimental trails on MuJoCo benchmarks demonstrate that the FP-DRL algorithm achieves state-of-the-art (SOTA) performance in most MuJoCo control tasks while exhibiting superior representation capability of the flow policy.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00977
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
Hao, Ruijie
Zhang, Longfei
Dai, Yang
Ma, Yang
Liang, Xingxing
Cheng, Guangquan
Machine Learning
Artificial Intelligence
Reinforcement Learning (RL) has proven highly effective in addressing complex control and decision-making tasks. However, in most traditional RL algorithms, the policy is typically parameterized as a diagonal Gaussian distribution, which constrains the policy from capturing multimodal distributions, making it difficult to cover the full range of optimal solutions in multi-solution problems, and the return is reduced to a mean value, losing its multimodal nature and thus providing insufficient guidance for policy updates. In response to these problems, we propose a RL algorithm termed flow-based policy with distributional RL (FP-DRL). This algorithm models the policy using flow matching, which offers both computational efficiency and the capacity to fit complex distributions. Additionally, it employs a distributional RL approach to model and optimize the entire return distribution, thereby more effectively guiding multimodal policy updates and improving agent performance. Experimental trails on MuJoCo benchmarks demonstrate that the FP-DRL algorithm achieves state-of-the-art (SOTA) performance in most MuJoCo control tasks while exhibiting superior representation capability of the flow policy.
title Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.00977