HybridFlow: A Two-Step Generative Policy for Robotic Manipulation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dong, Zhenchen, Fu, Jinna, Wu, Jiaming, Yu, Shengyuan, Chen, Fulin, Liu, Yide
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918338296283136
author Dong, Zhenchen
Fu, Jinna
Wu, Jiaming
Yu, Shengyuan
Chen, Fulin
Liu, Yide
author_facet Dong, Zhenchen
Fu, Jinna
Wu, Jiaming
Yu, Shengyuan
Chen, Fulin
Liu, Yide
contents Limited by inference latency, existing robot manipulation policies lack sufficient real-time interaction capability with the environment. Although faster generation methods such as flow matching are gradually replacing diffusion methods, researchers are pursuing even faster generation suitable for interactive robot control. MeanFlow, as a one-step variant of flow matching, has shown strong potential in image generation, but its precision in action generation does not meet the stringent requirements of robotic manipulation. We therefore propose \textbf{HybridFlow}, a \textbf{3-stage method} with \textbf{2-NFE}: Global Jump in MeanFlow mode, ReNoise for distribution alignment, and Local Refine in ReFlow mode. This method balances inference speed and generation quality by leveraging the rapid advantage of MeanFlow one-step generation while ensuring action precision with minimal generation steps. Through real-world experiments, HybridFlow outperforms the 16-step Diffusion Policy by \textbf{15--25\%} in success rate while reducing inference time from 152ms to 19ms (\textbf{8$\times$ speedup}, \textbf{$\sim$52Hz}); it also achieves 70.0\% success on unseen-color OOD grasping and 66.3\% on deformable object folding. We envision HybridFlow as a practical low-latency method to enhance real-world interaction capabilities of robotic manipulation policies.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13718
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HybridFlow: A Two-Step Generative Policy for Robotic Manipulation
Dong, Zhenchen
Fu, Jinna
Wu, Jiaming
Yu, Shengyuan
Chen, Fulin
Liu, Yide
Robotics
Artificial Intelligence
Limited by inference latency, existing robot manipulation policies lack sufficient real-time interaction capability with the environment. Although faster generation methods such as flow matching are gradually replacing diffusion methods, researchers are pursuing even faster generation suitable for interactive robot control. MeanFlow, as a one-step variant of flow matching, has shown strong potential in image generation, but its precision in action generation does not meet the stringent requirements of robotic manipulation. We therefore propose \textbf{HybridFlow}, a \textbf{3-stage method} with \textbf{2-NFE}: Global Jump in MeanFlow mode, ReNoise for distribution alignment, and Local Refine in ReFlow mode. This method balances inference speed and generation quality by leveraging the rapid advantage of MeanFlow one-step generation while ensuring action precision with minimal generation steps. Through real-world experiments, HybridFlow outperforms the 16-step Diffusion Policy by \textbf{15--25\%} in success rate while reducing inference time from 152ms to 19ms (\textbf{8$\times$ speedup}, \textbf{$\sim$52Hz}); it also achieves 70.0\% success on unseen-color OOD grasping and 66.3\% on deformable object folding. We envision HybridFlow as a practical low-latency method to enhance real-world interaction capabilities of robotic manipulation policies.
title HybridFlow: A Two-Step Generative Policy for Robotic Manipulation
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2602.13718