Pure Vision Language Action (VLA) Models: A Comprehensive Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Dapeng, Sun, Jing, Hu, Chenghui, Wu, Xiaoyan, Yuan, Zhenlong, Zhou, Rui, Shen, Fei, Zhou, Qingguo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915608261558272
author Zhang, Dapeng
Sun, Jing
Hu, Chenghui
Wu, Xiaoyan
Yuan, Zhenlong
Zhou, Rui
Shen, Fei
Zhou, Qingguo
author_facet Zhang, Dapeng
Sun, Jing
Hu, Chenghui
Wu, Xiaoyan
Yuan, Zhenlong
Zhou, Rui
Shen, Fei
Zhou, Qingguo
contents The emergence of Vision Language Action (VLA) models marks a paradigm shift from traditional policy-based control to generalized robotics, reframing Vision Language Models (VLMs) from passive sequence generators into active agents for manipulation and decision-making in complex, dynamic environments. This survey delves into advanced VLA methods, aiming to provide a clear taxonomy and a systematic, comprehensive review of existing research. It presents a comprehensive analysis of VLA applications across different scenarios and classifies VLA approaches into several paradigms: autoregression-based, diffusion-based, reinforcement-based, hybrid, and specialized methods; while examining their motivations, core strategies, and implementations in detail. In addition, foundational datasets, benchmarks, and simulation platforms are introduced. Building on the current VLA landscape, the review further proposes perspectives on key challenges and future directions to advance research in VLA models and generalizable robotics. By synthesizing insights from over three hundred recent studies, this survey maps the contours of this rapidly evolving field and highlights the opportunities and challenges that will shape the development of scalable, general-purpose VLA methods.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19012
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pure Vision Language Action (VLA) Models: A Comprehensive Survey
Zhang, Dapeng
Sun, Jing
Hu, Chenghui
Wu, Xiaoyan
Yuan, Zhenlong
Zhou, Rui
Shen, Fei
Zhou, Qingguo
Robotics
Artificial Intelligence
The emergence of Vision Language Action (VLA) models marks a paradigm shift from traditional policy-based control to generalized robotics, reframing Vision Language Models (VLMs) from passive sequence generators into active agents for manipulation and decision-making in complex, dynamic environments. This survey delves into advanced VLA methods, aiming to provide a clear taxonomy and a systematic, comprehensive review of existing research. It presents a comprehensive analysis of VLA applications across different scenarios and classifies VLA approaches into several paradigms: autoregression-based, diffusion-based, reinforcement-based, hybrid, and specialized methods; while examining their motivations, core strategies, and implementations in detail. In addition, foundational datasets, benchmarks, and simulation platforms are introduced. Building on the current VLA landscape, the review further proposes perspectives on key challenges and future directions to advance research in VLA models and generalizable robotics. By synthesizing insights from over three hundred recent studies, this survey maps the contours of this rapidly evolving field and highlights the opportunities and challenges that will shape the development of scalable, general-purpose VLA methods.
title Pure Vision Language Action (VLA) Models: A Comprehensive Survey
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2509.19012