A Survey on Efficient Vision-Language-Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Zhaoshu, Wang, Bo, Zeng, Pengpeng, Zhang, Haonan, Zhang, Ji, Wang, Zheng, Gao, Lianli, Song, Jingkuan, Sebe, Nicu, Shen, Heng Tao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908804112711680
author Yu, Zhaoshu
Wang, Bo
Zeng, Pengpeng
Zhang, Haonan
Zhang, Ji
Wang, Zheng
Gao, Lianli
Song, Jingkuan
Sebe, Nicu
Shen, Heng Tao
author_facet Yu, Zhaoshu
Wang, Bo
Zeng, Pengpeng
Zhang, Haonan
Zhang, Ji
Wang, Zheng
Gao, Lianli
Song, Jingkuan
Sebe, Nicu
Shen, Heng Tao
contents Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale architectures. While a surge of recent research has focused on enhancing VLA efficiency, the field lacks a unified framework to consolidate these disparate advancements. To bridge this gap, this survey presents the first comprehensive review of Efficient Vision-Language-Action models (Efficient VLAs) across the entire model-training-data pipeline. Specifically, we introduce a unified taxonomy to systematically organize the disparate efforts in this domain, categorizing current techniques into three core pillars: (1) Efficient Model Design, focusing on efficient architectures and model compression; (2) Efficient Training, which reduces computational burdens during model learning; and (3) Efficient Data Collection, which addresses the bottlenecks in acquiring and utilizing robotic data. Through a critical review of state-of-the-art methods within this framework, this survey not only establishes a foundational reference for the community but also summarizes representative applications, delineates key challenges, and charts a roadmap for future research. We maintain a continuously updated project page to track our latest developments: https://evla-survey.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24795
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey on Efficient Vision-Language-Action Models
Yu, Zhaoshu
Wang, Bo
Zeng, Pengpeng
Zhang, Haonan
Zhang, Ji
Wang, Zheng
Gao, Lianli
Song, Jingkuan
Sebe, Nicu
Shen, Heng Tao
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale architectures. While a surge of recent research has focused on enhancing VLA efficiency, the field lacks a unified framework to consolidate these disparate advancements. To bridge this gap, this survey presents the first comprehensive review of Efficient Vision-Language-Action models (Efficient VLAs) across the entire model-training-data pipeline. Specifically, we introduce a unified taxonomy to systematically organize the disparate efforts in this domain, categorizing current techniques into three core pillars: (1) Efficient Model Design, focusing on efficient architectures and model compression; (2) Efficient Training, which reduces computational burdens during model learning; and (3) Efficient Data Collection, which addresses the bottlenecks in acquiring and utilizing robotic data. Through a critical review of state-of-the-art methods within this framework, this survey not only establishes a foundational reference for the community but also summarizes representative applications, delineates key challenges, and charts a roadmap for future research. We maintain a continuously updated project page to track our latest developments: https://evla-survey.github.io/.
title A Survey on Efficient Vision-Language-Action Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2510.24795