Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Yuting, Ding, Leilei, Tang, Zhipeng, Zhu, Zenghuan, Deng, Jiajun, Lin, Xinrui, Liu, Shuo, Ren, Haojie, Ji, Jianmin, Zhang, Yanyong
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911413293809664
author Huang, Yuting
Ding, Leilei
Tang, Zhipeng
Zhu, Zenghuan
Deng, Jiajun
Lin, Xinrui
Liu, Shuo
Ren, Haojie
Ji, Jianmin
Zhang, Yanyong
author_facet Huang, Yuting
Ding, Leilei
Tang, Zhipeng
Zhu, Zenghuan
Deng, Jiajun
Lin, Xinrui
Liu, Shuo
Ren, Haojie
Ji, Jianmin
Zhang, Yanyong
contents While Vision-Language-Action (VLA) models hold promise in embodied intelligence, their large parameter counts lead to substantial inference latency that hinders real-time manipulation, motivating parameter sparsification. However, as the environment evolves during VLA execution, the optimal sparsity patterns change accordingly. Static pruning lacks the adaptability required for environment dynamics, whereas fixed-interval dynamic layer pruning suffers from coarse granularity and high retraining overheads. To bridge this gap, we propose EcoVLA, a training-free, plug-and-play adaptive pruning framework that supports orthogonal combination with existing VLA acceleration methods. EcoVLA comprises two components: Environment-aware Adaptive Pruning (EAP) and Interleaved Inference Orchestration ($I^2O$). EAP is a lightweight adaptive channel pruning method that incorporates the temporal consistency of the physical environment to update sparsity patterns. $I^2O$ leverages the FLOPs bubbles inherent in VLA inference to schedule the pruning method in parallel, ensuring negligible impact on latency. Evaluated on diverse VLA models and benchmarks, EcoVLA delivers state-of-the-art performance, achieving up to 1.60$\times$ speedup with only a 0.4% drop in success rate, and further reaches 2.18$\times$ speedup with only a 0.5% degradation when combined with token pruning. We further validate the effectiveness of EcoVLA on real-world robots.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00780
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models
Huang, Yuting
Ding, Leilei
Tang, Zhipeng
Zhu, Zenghuan
Deng, Jiajun
Lin, Xinrui
Liu, Shuo
Ren, Haojie
Ji, Jianmin
Zhang, Yanyong
Artificial Intelligence
While Vision-Language-Action (VLA) models hold promise in embodied intelligence, their large parameter counts lead to substantial inference latency that hinders real-time manipulation, motivating parameter sparsification. However, as the environment evolves during VLA execution, the optimal sparsity patterns change accordingly. Static pruning lacks the adaptability required for environment dynamics, whereas fixed-interval dynamic layer pruning suffers from coarse granularity and high retraining overheads. To bridge this gap, we propose EcoVLA, a training-free, plug-and-play adaptive pruning framework that supports orthogonal combination with existing VLA acceleration methods. EcoVLA comprises two components: Environment-aware Adaptive Pruning (EAP) and Interleaved Inference Orchestration ($I^2O$). EAP is a lightweight adaptive channel pruning method that incorporates the temporal consistency of the physical environment to update sparsity patterns. $I^2O$ leverages the FLOPs bubbles inherent in VLA inference to schedule the pruning method in parallel, ensuring negligible impact on latency. Evaluated on diverse VLA models and benchmarks, EcoVLA delivers state-of-the-art performance, achieving up to 1.60$\times$ speedup with only a 0.4% drop in success rate, and further reaches 2.18$\times$ speedup with only a 0.5% degradation when combined with token pruning. We further validate the effectiveness of EcoVLA on real-world robots.
title Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models
topic Artificial Intelligence
url https://arxiv.org/abs/2602.00780