ChainFlow-VLA: Causal Flow Planning with Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiyang, Wang, Xinlin, Zhou, Tingguang, Chen, Gong, Gui, Xingtai, Xu, Zhi, Wu, Xiaolei, Tan, Feiyang, Zhou, Hangning, Yang, Mu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914590974017536
author Wang, Xiyang
Wang, Xinlin
Zhou, Tingguang
Chen, Gong
Gui, Xingtai
Xu, Zhi
Wu, Xiaolei
Tan, Feiyang
Zhou, Hangning
Yang, Mu
author_facet Wang, Xiyang
Wang, Xinlin
Zhou, Tingguang
Chen, Gong
Gui, Xingtai
Xu, Zhi
Wu, Xiaolei
Tan, Feiyang
Zhou, Hangning
Yang, Mu
contents Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction-aware temporal dependencies via causal factorization, but their step-wise decoding leads to error accumulation and suboptimal global structure. In contrast, diffusion models optimize trajectories globally but lack explicit causal constraints, making them unreliable in interactive and safety-critical scenarios. This dichotomy reveals a deeper issue: existing methods treat causal modeling and global optimization as separate paradigms, without a principled way to unify them within a single trajectory distribution. To address this, we propose ChainFlow-VLA, which unifies causal generation and global refinement within a unified probabilistic framework. We formulate planning as a mixture over AR-induced modes and learn Vision-Language Model (VLM)-conditioned residual distributions over these modes. An autoregressive generator (Chain) produces a discrete set of causal trajectory modes, followed by a diffusion-based refiner (Flow) that leverages VLM hidden states as semantic priors to perform mode-conditioned correction in residual space while preserving causal structure. This straightforward conditioning seamlessly injects high-level scene understanding into fine-grained trajectory adjustments. Experiments demonstrate that ChainFlow-VLA achieves robust planning in ambiguous and long-tail scenarios, achieving a state-of-the-art score of 94.85 on the NAVSIM v1 leaderboard, matching human-level performance (94.8). Code will be available at https://github.com/AFARI-Research/ChainFlow-VLA.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23270
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ChainFlow-VLA: Causal Flow Planning with Vision-Language Models
Wang, Xiyang
Wang, Xinlin
Zhou, Tingguang
Chen, Gong
Gui, Xingtai
Xu, Zhi
Wu, Xiaolei
Tan, Feiyang
Zhou, Hangning
Yang, Mu
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction-aware temporal dependencies via causal factorization, but their step-wise decoding leads to error accumulation and suboptimal global structure. In contrast, diffusion models optimize trajectories globally but lack explicit causal constraints, making them unreliable in interactive and safety-critical scenarios. This dichotomy reveals a deeper issue: existing methods treat causal modeling and global optimization as separate paradigms, without a principled way to unify them within a single trajectory distribution. To address this, we propose ChainFlow-VLA, which unifies causal generation and global refinement within a unified probabilistic framework. We formulate planning as a mixture over AR-induced modes and learn Vision-Language Model (VLM)-conditioned residual distributions over these modes. An autoregressive generator (Chain) produces a discrete set of causal trajectory modes, followed by a diffusion-based refiner (Flow) that leverages VLM hidden states as semantic priors to perform mode-conditioned correction in residual space while preserving causal structure. This straightforward conditioning seamlessly injects high-level scene understanding into fine-grained trajectory adjustments. Experiments demonstrate that ChainFlow-VLA achieves robust planning in ambiguous and long-tail scenarios, achieving a state-of-the-art score of 94.85 on the NAVSIM v1 leaderboard, matching human-level performance (94.8). Code will be available at https://github.com/AFARI-Research/ChainFlow-VLA.
title ChainFlow-VLA: Causal Flow Planning with Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2605.23270