FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhong, Zhide, Yan, Haodong, Li, Junfeng, Liu, Xiangchen, Gong, Xin, Zhang, Tianran, Song, Wenxuan, Chen, Jiayi, Zheng, Xinhu, Wang, Hesheng, Li, Haoang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912632953372672
author Zhong, Zhide
Yan, Haodong
Li, Junfeng
Liu, Xiangchen
Gong, Xin
Zhang, Tianran
Song, Wenxuan
Chen, Jiayi
Zheng, Xinhu
Wang, Hesheng
Li, Haoang
author_facet Zhong, Zhide
Yan, Haodong
Li, Junfeng
Liu, Xiangchen
Gong, Xin
Zhang, Tianran
Song, Wenxuan
Chen, Jiayi
Zheng, Xinhu
Wang, Hesheng
Li, Haoang
contents Many Vision-Language-Action (VLA) models are built upon an internal world model trained via next-frame prediction ``$v_t \rightarrow v_{t+1}$''. However, this paradigm attempts to predict the future frame's appearance directly, without explicitly reasoning about the underlying dynamics. \textbf{This lack of an explicit motion reasoning step} often leads to physically implausible visual forecasts and inefficient policy learning. To address this limitation, we introduce the \textbf{Visual Chain of Thought (Visual CoT)}, a paradigm that compels the model to first reason about \textbf{motion dynamics} before generating the future frame. We instantiate this paradigm by proposing \textbf{FlowVLA}, an autoregressive Transformer that explicitly materializes this reasoning process as ``$v_t \rightarrow f_t \rightarrow v_{t+1}$'', where $f_t$ is an intermediate optical flow prediction that inherently encodes motion. By forcing the model to first follow the motion plan encoded by $f_t$, this process inherently \textbf{aligns the pre-training objective of dynamics prediction with the downstream task of action generation.} We conduct experiments on challenging robotics manipulation benchmarks, as well as real-robot evaluations. Our FlowVLA not only generates \textbf{more coherent and physically plausible visual predictions}, but also achieves state-of-the-art policy performance with \textbf{substantially improved sample efficiency}, pointing toward a more principled foundation for world modeling in VLAs. Project page: https://irpn-lab.github.io/FlowVLA/
format Preprint
id arxiv_https___arxiv_org_abs_2508_18269
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
Zhong, Zhide
Yan, Haodong
Li, Junfeng
Liu, Xiangchen
Gong, Xin
Zhang, Tianran
Song, Wenxuan
Chen, Jiayi
Zheng, Xinhu
Wang, Hesheng
Li, Haoang
Robotics
Many Vision-Language-Action (VLA) models are built upon an internal world model trained via next-frame prediction ``$v_t \rightarrow v_{t+1}$''. However, this paradigm attempts to predict the future frame's appearance directly, without explicitly reasoning about the underlying dynamics. \textbf{This lack of an explicit motion reasoning step} often leads to physically implausible visual forecasts and inefficient policy learning. To address this limitation, we introduce the \textbf{Visual Chain of Thought (Visual CoT)}, a paradigm that compels the model to first reason about \textbf{motion dynamics} before generating the future frame. We instantiate this paradigm by proposing \textbf{FlowVLA}, an autoregressive Transformer that explicitly materializes this reasoning process as ``$v_t \rightarrow f_t \rightarrow v_{t+1}$'', where $f_t$ is an intermediate optical flow prediction that inherently encodes motion. By forcing the model to first follow the motion plan encoded by $f_t$, this process inherently \textbf{aligns the pre-training objective of dynamics prediction with the downstream task of action generation.} We conduct experiments on challenging robotics manipulation benchmarks, as well as real-robot evaluations. Our FlowVLA not only generates \textbf{more coherent and physically plausible visual predictions}, but also achieves state-of-the-art policy performance with \textbf{substantially improved sample efficiency}, pointing toward a more principled foundation for world modeling in VLAs. Project page: https://irpn-lab.github.io/FlowVLA/
title FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
topic Robotics
url https://arxiv.org/abs/2508.18269