SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yixian, Yu, Shu'ang, Zhang, Tonghe, Guang, Mo, Hui, Haojia, Long, Kaiwen, Wang, Yu, Yu, Chao, Ding, Wenbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915729064853504
author Zhang, Yixian
Yu, Shu'ang
Zhang, Tonghe
Guang, Mo
Hui, Haojia
Long, Kaiwen
Wang, Yu
Yu, Chao
Ding, Wenbo
author_facet Zhang, Yixian
Yu, Shu'ang
Zhang, Tonghe
Guang, Mo
Hui, Haojia
Long, Kaiwen
Wang, Yu
Yu, Chao
Ding, Wenbo
contents Training expressive flow-based policies with off-policy reinforcement learning is notoriously unstable due to gradient pathologies in the multi-step action sampling process. We trace this instability to a fundamental connection: the flow rollout is algebraically equivalent to a residual recurrent computation, making it susceptible to the same vanishing and exploding gradients as RNNs. To address this, we reparameterize the velocity network using principles from modern sequential models, introducing two stable architectures: Flow-G, which incorporates a gated velocity, and Flow-T, which utilizes a decoded velocity. We then develop a practical SAC-based algorithm, enabled by a noise-augmented rollout, that facilitates direct end-to-end training of these policies. Our approach supports both from-scratch and offline-to-online learning and achieves state-of-the-art performance on continuous control and robotic manipulation benchmarks, eliminating the need for common workarounds like policy distillation or surrogate objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25756
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling
Zhang, Yixian
Yu, Shu'ang
Zhang, Tonghe
Guang, Mo
Hui, Haojia
Long, Kaiwen
Wang, Yu
Yu, Chao
Ding, Wenbo
Robotics
Machine Learning
Training expressive flow-based policies with off-policy reinforcement learning is notoriously unstable due to gradient pathologies in the multi-step action sampling process. We trace this instability to a fundamental connection: the flow rollout is algebraically equivalent to a residual recurrent computation, making it susceptible to the same vanishing and exploding gradients as RNNs. To address this, we reparameterize the velocity network using principles from modern sequential models, introducing two stable architectures: Flow-G, which incorporates a gated velocity, and Flow-T, which utilizes a decoded velocity. We then develop a practical SAC-based algorithm, enabled by a noise-augmented rollout, that facilitates direct end-to-end training of these policies. Our approach supports both from-scratch and offline-to-online learning and achieves state-of-the-art performance on continuous control and robotic manipulation benchmarks, eliminating the need for common workarounds like policy distillation or surrogate objectives.
title SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling
topic Robotics
Machine Learning
url https://arxiv.org/abs/2509.25756