Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.14049 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908455713898496 |
|---|---|
| author | Budzianowski, Paweł Maa, Wesley Freed, Matthew Mo, Jingxiang Hsiao, Winston Xie, Aaron Młoduchowski, Tomasz Tipnis, Viraj Bolte, Benjamin |
| author_facet | Budzianowski, Paweł Maa, Wesley Freed, Matthew Mo, Jingxiang Hsiao, Winston Xie, Aaron Młoduchowski, Tomasz Tipnis, Viraj Bolte, Benjamin |
| contents | Vision-Language Models (VLMs) have emerged as a promising approach to address the data scarcity challenge in robotics, enabling the development of generalizable visuomotor control policies. While models like OpenVLA showcase the potential of this paradigm, deploying large-scale VLMs on resource-constrained mobile manipulation systems remains a significant hurdle. This paper introduces Edge VLA (EVLA), a novel approach designed to significantly enhance the inference speed of Vision-Language-Action (VLA) models. EVLA maintains the representational power of these models while enabling real-time performance on edge devices. We achieve this through two key innovations: 1) Eliminating the autoregressive requirement for end-effector position prediction, leading to a 7x speedup in inference, and 2) Leveraging the efficiency of Small Language Models (SLMs), demonstrating comparable training performance to larger models with significantly reduced computational demands. Our early results demonstrate that EVLA achieves comparable training characteristics to OpenVLA while offering substantial gains in inference speed and memory efficiency. We release our model checkpoints and training \href{https://github.com/kscalelabs/evla }{codebase} to foster further research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_14049 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | EdgeVLA: Efficient Vision-Language-Action Models Budzianowski, Paweł Maa, Wesley Freed, Matthew Mo, Jingxiang Hsiao, Winston Xie, Aaron Młoduchowski, Tomasz Tipnis, Viraj Bolte, Benjamin Robotics Computation and Language Vision-Language Models (VLMs) have emerged as a promising approach to address the data scarcity challenge in robotics, enabling the development of generalizable visuomotor control policies. While models like OpenVLA showcase the potential of this paradigm, deploying large-scale VLMs on resource-constrained mobile manipulation systems remains a significant hurdle. This paper introduces Edge VLA (EVLA), a novel approach designed to significantly enhance the inference speed of Vision-Language-Action (VLA) models. EVLA maintains the representational power of these models while enabling real-time performance on edge devices. We achieve this through two key innovations: 1) Eliminating the autoregressive requirement for end-effector position prediction, leading to a 7x speedup in inference, and 2) Leveraging the efficiency of Small Language Models (SLMs), demonstrating comparable training performance to larger models with significantly reduced computational demands. Our early results demonstrate that EVLA achieves comparable training characteristics to OpenVLA while offering substantial gains in inference speed and memory efficiency. We release our model checkpoints and training \href{https://github.com/kscalelabs/evla }{codebase} to foster further research. |
| title | EdgeVLA: Efficient Vision-Language-Action Models |
| topic | Robotics Computation and Language |
| url | https://arxiv.org/abs/2507.14049 |