SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Ye, Meng, Yuan, Sun, Zewen, Ji, Kangye, Tang, Chen, Fan, Jiajun, Ma, Xinzhu, Xia, Shutao, Wang, Zhi, Zhu, Wenwu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914072274927616
author Li, Ye
Meng, Yuan
Sun, Zewen
Ji, Kangye
Tang, Chen
Fan, Jiajun
Ma, Xinzhu
Xia, Shutao
Wang, Zhi
Zhu, Wenwu
author_facet Li, Ye
Meng, Yuan
Sun, Zewen
Ji, Kangye
Tang, Chen
Fan, Jiajun
Ma, Xinzhu
Xia, Shutao
Wang, Zhi
Zhu, Wenwu
contents Vision-Language-Action (VLA) models have attracted increasing attention for their strong control capabilities. However, their high computational cost and low execution frequency hinder their suitability for real-time tasks such as robotic manipulation and autonomous navigation. Existing VLA acceleration methods primarily focus on structural optimization, overlooking the fact that these models operate in sequential decision-making environments. As a result, temporal redundancy in sequential action generation and spatial redundancy in visual input remain unaddressed. To this end, we propose SP-VLA, a unified framework that accelerates VLA models by jointly scheduling models and pruning tokens. Specifically, we design an action-aware model scheduling mechanism that reduces temporal redundancy by dynamically switching between VLA model and a lightweight generator. Inspired by the human motion pattern of focusing on key decision points while relying on intuition for other actions, we categorize VLA actions into deliberative and intuitive, assigning the former to the VLA model and the latter to the lightweight generator, enabling frequency-adaptive execution through collaborative model scheduling. To address spatial redundancy, we further develop a spatio-semantic dual-aware token pruning method. Tokens are classified into spatial and semantic types and pruned based on their dual-aware importance to accelerate VLA inference. These two mechanisms work jointly to guide the VLA in focusing on critical actions and salient visual information, achieving effective acceleration while maintaining high accuracy. Extensive experiments show that our method achieves 1.5$\times$ lossless acceleration in LIBERO and 2.4$\times$ in SimplerEnv, with up to 6% average performance gain. Inference frequency and latency improve by 2.2$\times$ in SimplerEnv and 1.4$\times$ in LIBERO.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12723
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration
Li, Ye
Meng, Yuan
Sun, Zewen
Ji, Kangye
Tang, Chen
Fan, Jiajun
Ma, Xinzhu
Xia, Shutao
Wang, Zhi
Zhu, Wenwu
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-Language-Action (VLA) models have attracted increasing attention for their strong control capabilities. However, their high computational cost and low execution frequency hinder their suitability for real-time tasks such as robotic manipulation and autonomous navigation. Existing VLA acceleration methods primarily focus on structural optimization, overlooking the fact that these models operate in sequential decision-making environments. As a result, temporal redundancy in sequential action generation and spatial redundancy in visual input remain unaddressed. To this end, we propose SP-VLA, a unified framework that accelerates VLA models by jointly scheduling models and pruning tokens. Specifically, we design an action-aware model scheduling mechanism that reduces temporal redundancy by dynamically switching between VLA model and a lightweight generator. Inspired by the human motion pattern of focusing on key decision points while relying on intuition for other actions, we categorize VLA actions into deliberative and intuitive, assigning the former to the VLA model and the latter to the lightweight generator, enabling frequency-adaptive execution through collaborative model scheduling. To address spatial redundancy, we further develop a spatio-semantic dual-aware token pruning method. Tokens are classified into spatial and semantic types and pruned based on their dual-aware importance to accelerate VLA inference. These two mechanisms work jointly to guide the VLA in focusing on critical actions and salient visual information, achieving effective acceleration while maintaining high accuracy. Extensive experiments show that our method achieves 1.5$\times$ lossless acceleration in LIBERO and 2.4$\times$ in SimplerEnv, with up to 6% average performance gain. Inference frequency and latency improve by 2.2$\times$ in SimplerEnv and 1.4$\times$ in LIBERO.
title SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.12723