SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Shilin, Zhang, Chubin, Wang, Changyuan, Wang, Yuji, Wu, Yue, Wang, Zixuan, Tian, Jingqi, Zhu, Zheng, Tang, Yansong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914613134622720
author Ma, Shilin
Zhang, Chubin
Wang, Changyuan
Wang, Yuji
Wu, Yue
Wang, Zixuan
Tian, Jingqi
Zhu, Zheng
Tang, Yansong
author_facet Ma, Shilin
Zhang, Chubin
Wang, Changyuan
Wang, Yuji
Wu, Yue
Wang, Zixuan
Tian, Jingqi
Zhu, Zheng
Tang, Yansong
contents Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most existing methods mainly base pruning decisions on shallow-layer cues and risk discarding visual information required by deep layers. To address this issue, we propose SAFE-Pruner, a plug-and-play pruning framework that incorporates attention cues of future layers into pruning decisions. Specifically, we identify semantic attention consistency, the tendency that VLA models concentrate their attention probability mass on the same semantic entity across execution steps. Based on this observation, we design a forward-looking strategy to forecast the token saliency in deep layers, which prevents the premature removal of critical tokens and leads to more stable acceleration. We further introduce an adaptive subtask division strategy to detect abrupt attention shifts, thereby improving forecasting accuracy and pruning reliability. Extensive experiments in simulation and real-world settings demonstrate that our method achieves up to 1.89x speedup with a minimal degradation in success rate of less than 1.7%, while outperforming state-of-the-art methods by up to 1.9%.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29662
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
Ma, Shilin
Zhang, Chubin
Wang, Changyuan
Wang, Yuji
Wu, Yue
Wang, Zixuan
Tian, Jingqi
Zhu, Zheng
Tang, Yansong
Computer Vision and Pattern Recognition
Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most existing methods mainly base pruning decisions on shallow-layer cues and risk discarding visual information required by deep layers. To address this issue, we propose SAFE-Pruner, a plug-and-play pruning framework that incorporates attention cues of future layers into pruning decisions. Specifically, we identify semantic attention consistency, the tendency that VLA models concentrate their attention probability mass on the same semantic entity across execution steps. Based on this observation, we design a forward-looking strategy to forecast the token saliency in deep layers, which prevents the premature removal of critical tokens and leads to more stable acceleration. We further introduce an adaptive subtask division strategy to detect abrupt attention shifts, thereby improving forecasting accuracy and pruning reliability. Extensive experiments in simulation and real-world settings demonstrate that our method achieves up to 1.89x speedup with a minimal degradation in success rate of less than 1.7%, while outperforming state-of-the-art methods by up to 1.9%.
title SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.29662