EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yantai, Wang, Yuhao, Wen, Zichen, Zhongwei, Luo, Zou, Chang, Zhang, Zhipeng, Wen, Chuan, Zhang, Linfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913889562656768
author Yang, Yantai
Wang, Yuhao
Wen, Zichen
Zhongwei, Luo
Zou, Chang
Zhang, Zhipeng
Wen, Chuan
Zhang, Linfeng
author_facet Yang, Yantai
Wang, Yuhao
Wen, Zichen
Zhongwei, Luo
Zou, Chang
Zhang, Zhipeng
Wen, Chuan
Zhang, Linfeng
contents Vision-Language-Action (VLA) models, particularly diffusion-based architectures, demonstrate transformative potential for embodied intelligence but are severely hampered by high computational and memory demands stemming from extensive inherent and inference-time redundancies. While existing acceleration efforts often target isolated inefficiencies, such piecemeal solutions typically fail to holistically address the varied computational and memory bottlenecks across the entire VLA pipeline, thereby limiting practical deployability. We introduce EfficientVLA, a structured and training-free inference acceleration framework that systematically eliminates these barriers by cohesively exploiting multifaceted redundancies. EfficientVLA synergistically integrates three targeted strategies: (1) pruning of functionally inconsequential layers from the language module, guided by an analysis of inter-layer redundancies; (2) optimizing the visual processing pathway through a task-aware strategy that selects a compact, diverse set of visual tokens, balancing task-criticality with informational coverage; and (3) alleviating temporal computational redundancy within the iterative diffusion-based action head by strategically caching and reusing key intermediate features. We apply our method to a standard VLA model CogACT, yielding a 1.93X inference speedup and reduces FLOPs to 28.9%, with only a 0.6% success rate drop in the SIMPLER benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10100
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
Yang, Yantai
Wang, Yuhao
Wen, Zichen
Zhongwei, Luo
Zou, Chang
Zhang, Zhipeng
Wen, Chuan
Zhang, Linfeng
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models, particularly diffusion-based architectures, demonstrate transformative potential for embodied intelligence but are severely hampered by high computational and memory demands stemming from extensive inherent and inference-time redundancies. While existing acceleration efforts often target isolated inefficiencies, such piecemeal solutions typically fail to holistically address the varied computational and memory bottlenecks across the entire VLA pipeline, thereby limiting practical deployability. We introduce EfficientVLA, a structured and training-free inference acceleration framework that systematically eliminates these barriers by cohesively exploiting multifaceted redundancies. EfficientVLA synergistically integrates three targeted strategies: (1) pruning of functionally inconsequential layers from the language module, guided by an analysis of inter-layer redundancies; (2) optimizing the visual processing pathway through a task-aware strategy that selects a compact, diverse set of visual tokens, balancing task-criticality with informational coverage; and (3) alleviating temporal computational redundancy within the iterative diffusion-based action head by strategically caching and reusing key intermediate features. We apply our method to a standard VLA model CogACT, yielding a 1.93X inference speedup and reduces FLOPs to 28.9%, with only a 0.6% success rate drop in the SIMPLER benchmark.
title EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.10100