AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Yuhua, Cheng, Shuang, Ding, Yan, Gao, Feifei, Qi, Biqing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913097832202240
author Jiang, Yuhua
Cheng, Shuang
Ding, Yan
Gao, Feifei
Qi, Biqing
author_facet Jiang, Yuhua
Cheng, Shuang
Ding, Yan
Gao, Feifei
Qi, Biqing
contents Vision-language-action (VLA) models have recently emerged as a powerful paradigm for building generalist robots. However, traditional VLA models that generate actions through flow matching (FM) typically rely on rigid and uniform time schedules, i.e., synchronous FM (SFM). Without action context awareness and asynchronous self-correction, SFM becomes unstable in long-horizon tasks, where a single action error can cascade into failure. In this work, we propose asynchronous flow matching VLA (AsyncVLA), a novel framework that introduces temporal flexibility in asynchronous FM (AFM) and enables self-correction in action generation. AsyncVLA breaks from the vanilla SFM in VLA models by generating the action tokens in a non-uniform time schedule with action context awareness. Besides, our method introduces the confidence rater to extract confidence of the initially generated actions, enabling the model to selectively refine inaccurate action tokens before execution. Moreover, we propose a unified training procedure for SFM and AFM that endows a single model with both modes, improving KV-cache utilization. Extensive experiments on robotic manipulation benchmarks demonstrate that AsyncVLA is data-efficient and exhibits self-correction ability. AsyncVLA outperforms existing methods across both simulation and real-world evaluations. Our code is available at https://github.com/YuhuaJiang2002/AsyncVLA.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14148
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
Jiang, Yuhua
Cheng, Shuang
Ding, Yan
Gao, Feifei
Qi, Biqing
Robotics
Artificial Intelligence
Machine Learning
Vision-language-action (VLA) models have recently emerged as a powerful paradigm for building generalist robots. However, traditional VLA models that generate actions through flow matching (FM) typically rely on rigid and uniform time schedules, i.e., synchronous FM (SFM). Without action context awareness and asynchronous self-correction, SFM becomes unstable in long-horizon tasks, where a single action error can cascade into failure. In this work, we propose asynchronous flow matching VLA (AsyncVLA), a novel framework that introduces temporal flexibility in asynchronous FM (AFM) and enables self-correction in action generation. AsyncVLA breaks from the vanilla SFM in VLA models by generating the action tokens in a non-uniform time schedule with action context awareness. Besides, our method introduces the confidence rater to extract confidence of the initially generated actions, enabling the model to selectively refine inaccurate action tokens before execution. Moreover, we propose a unified training procedure for SFM and AFM that endows a single model with both modes, improving KV-cache utilization. Extensive experiments on robotic manipulation benchmarks demonstrate that AsyncVLA is data-efficient and exhibits self-correction ability. AsyncVLA outperforms existing methods across both simulation and real-world evaluations. Our code is available at https://github.com/YuhuaJiang2002/AsyncVLA.
title AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.14148