A Survey on Efficient Vision-Language-Action Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908804112711680 |
|---|---|
| author | Yu, Zhaoshu Wang, Bo Zeng, Pengpeng Zhang, Haonan Zhang, Ji Wang, Zheng Gao, Lianli Song, Jingkuan Sebe, Nicu Shen, Heng Tao |
| author_facet | Yu, Zhaoshu Wang, Bo Zeng, Pengpeng Zhang, Haonan Zhang, Ji Wang, Zheng Gao, Lianli Song, Jingkuan Sebe, Nicu Shen, Heng Tao |
| contents | Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale architectures. While a surge of recent research has focused on enhancing VLA efficiency, the field lacks a unified framework to consolidate these disparate advancements. To bridge this gap, this survey presents the first comprehensive review of Efficient Vision-Language-Action models (Efficient VLAs) across the entire model-training-data pipeline. Specifically, we introduce a unified taxonomy to systematically organize the disparate efforts in this domain, categorizing current techniques into three core pillars: (1) Efficient Model Design, focusing on efficient architectures and model compression; (2) Efficient Training, which reduces computational burdens during model learning; and (3) Efficient Data Collection, which addresses the bottlenecks in acquiring and utilizing robotic data. Through a critical review of state-of-the-art methods within this framework, this survey not only establishes a foundational reference for the community but also summarizes representative applications, delineates key challenges, and charts a roadmap for future research. We maintain a continuously updated project page to track our latest developments: https://evla-survey.github.io/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_24795 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Survey on Efficient Vision-Language-Action Models Yu, Zhaoshu Wang, Bo Zeng, Pengpeng Zhang, Haonan Zhang, Ji Wang, Zheng Gao, Lianli Song, Jingkuan Sebe, Nicu Shen, Heng Tao Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Robotics Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale architectures. While a surge of recent research has focused on enhancing VLA efficiency, the field lacks a unified framework to consolidate these disparate advancements. To bridge this gap, this survey presents the first comprehensive review of Efficient Vision-Language-Action models (Efficient VLAs) across the entire model-training-data pipeline. Specifically, we introduce a unified taxonomy to systematically organize the disparate efforts in this domain, categorizing current techniques into three core pillars: (1) Efficient Model Design, focusing on efficient architectures and model compression; (2) Efficient Training, which reduces computational burdens during model learning; and (3) Efficient Data Collection, which addresses the bottlenecks in acquiring and utilizing robotic data. Through a critical review of state-of-the-art methods within this framework, this survey not only establishes a foundational reference for the community but also summarizes representative applications, delineates key challenges, and charts a roadmap for future research. We maintain a continuously updated project page to track our latest developments: https://evla-survey.github.io/. |
| title | A Survey on Efficient Vision-Language-Action Models |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Robotics |
| url | https://arxiv.org/abs/2510.24795 |