A Survey on Vision-Language-Action Models for Embodied AI

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ma, Yueen, Song, Zixing, Zhuang, Yuzheng, Hao, Jianye, King, Irwin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913078126313472
author Ma, Yueen
Song, Zixing
Zhuang, Yuzheng
Hao, Jianye
King, Irwin
author_facet Ma, Yueen
Song, Zixing
Zhuang, Yuzheng
Hao, Jianye
King, Irwin
contents Embodied AI is widely recognized as a cornerstone of artificial general intelligence (AGI) because it involves controlling embodied agents to perform tasks in the physical world. Building on the success of large language models (LLMs) and vision-language models (VLMs), a new category of multimodal models -- referred to as vision-language-action (VLA) models -- has emerged to address language-conditioned robotic tasks in embodied AI by leveraging their distinct ability to generate actions. The recent proliferation of VLAs necessitates a comprehensive survey to capture the rapidly evolving landscape. To this end, we present the first survey on VLAs for embodied AI. This work provides a detailed taxonomy of VLAs, organized into three major lines of research. The first line focuses on individual components of VLAs. The second line is dedicated to developing VLA-based control policies adept at predicting low-level actions. The third line comprises high-level task planners capable of decomposing long-horizon tasks into a sequence of subtasks, thereby guiding VLAs to follow more general user instructions. Furthermore, we provide an extensive summary of relevant resources, including datasets, simulators, and benchmarks. Finally, we discuss the challenges facing VLAs and outline promising future directions in embodied AI. A curated repository associated with this survey is available at: https://github.com/yueen-ma/Awesome-VLA.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14093
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Survey on Vision-Language-Action Models for Embodied AI
Ma, Yueen
Song, Zixing
Zhuang, Yuzheng
Hao, Jianye
King, Irwin
Robotics
Computation and Language
Computer Vision and Pattern Recognition
68T40 (Primary) 68T45, 68T50, 68T07 (Secondary)
I.2.9; I.2.10; I.2.7; I.2.6
Embodied AI is widely recognized as a cornerstone of artificial general intelligence (AGI) because it involves controlling embodied agents to perform tasks in the physical world. Building on the success of large language models (LLMs) and vision-language models (VLMs), a new category of multimodal models -- referred to as vision-language-action (VLA) models -- has emerged to address language-conditioned robotic tasks in embodied AI by leveraging their distinct ability to generate actions. The recent proliferation of VLAs necessitates a comprehensive survey to capture the rapidly evolving landscape. To this end, we present the first survey on VLAs for embodied AI. This work provides a detailed taxonomy of VLAs, organized into three major lines of research. The first line focuses on individual components of VLAs. The second line is dedicated to developing VLA-based control policies adept at predicting low-level actions. The third line comprises high-level task planners capable of decomposing long-horizon tasks into a sequence of subtasks, thereby guiding VLAs to follow more general user instructions. Furthermore, we provide an extensive summary of relevant resources, including datasets, simulators, and benchmarks. Finally, we discuss the challenges facing VLAs and outline promising future directions in embodied AI. A curated repository associated with this survey is available at: https://github.com/yueen-ma/Awesome-VLA.
title A Survey on Vision-Language-Action Models for Embodied AI
topic Robotics
Computation and Language
Computer Vision and Pattern Recognition
68T40 (Primary) 68T45, 68T50, 68T07 (Secondary)
I.2.9; I.2.10; I.2.7; I.2.6
url https://arxiv.org/abs/2405.14093