Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Yulin, Chen, Hao, Wu, Zhuangzhe, Sui, Bowen, Liu, Jiaming, Gu, Chenyang, Liu, Zhuoyang, Feng, Qiuxuan, Yu, Jiale, Gu, Shuo, Jia, Peng, Heng, Pheng-Ann, Zhang, Shanghang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911521656799232
author Luo, Yulin
Chen, Hao
Wu, Zhuangzhe
Sui, Bowen
Liu, Jiaming
Gu, Chenyang
Liu, Zhuoyang
Feng, Qiuxuan
Yu, Jiale
Gu, Shuo
Jia, Peng
Heng, Pheng-Ann
Zhang, Shanghang
author_facet Luo, Yulin
Chen, Hao
Wu, Zhuangzhe
Sui, Bowen
Liu, Jiaming
Gu, Chenyang
Liu, Zhuoyang
Feng, Qiuxuan
Yu, Jiale
Gu, Shuo
Jia, Peng
Heng, Pheng-Ann
Zhang, Shanghang
contents Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for robotic manipulation, in which reliable action prediction critically depends on accurately interpreting and integrating visual observations conditioned on language instructions. Although recent works have sought to enhance the visual capabilities of VLA models, most approaches treat the LLM backbone as a black box, providing limited insight into how visual information is grounded into action generation. Therefore, we perform a systematic analysis of multiple VLA models across different action-generation paradigms and observe that sensitivity to visual tokens progressively decreases in deeper layers during action generation. Motivated by this observation, we propose \textbf{DeepVision-VLA}, built on a \textbf{Vision-Language Mixture-of-Transformers (VL-MoT)} framework. This framework enables shared attention between the vision foundation model and the VLA backbone, injecting multi-level visual features from the vision expert into deeper layers of the VLA backbone to enhance visual representations for precise and complex manipulation. In addition, we introduce \textbf{Action-Guided Visual Pruning (AGVP)}, which leverages shallow-layer attention to prune irrelevant visual tokens while preserving task-relevant ones, reinforcing critical visual cues for manipulation with minimal computational overhead. DeepVision-VLA outperforms prior state-of-the-art methods by 9.0\% and 7.5\% on simulated and real-world tasks, respectively, providing new insights for the design of visually enhanced VLA models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15618
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
Luo, Yulin
Chen, Hao
Wu, Zhuangzhe
Sui, Bowen
Liu, Jiaming
Gu, Chenyang
Liu, Zhuoyang
Feng, Qiuxuan
Yu, Jiale
Gu, Shuo
Jia, Peng
Heng, Pheng-Ann
Zhang, Shanghang
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for robotic manipulation, in which reliable action prediction critically depends on accurately interpreting and integrating visual observations conditioned on language instructions. Although recent works have sought to enhance the visual capabilities of VLA models, most approaches treat the LLM backbone as a black box, providing limited insight into how visual information is grounded into action generation. Therefore, we perform a systematic analysis of multiple VLA models across different action-generation paradigms and observe that sensitivity to visual tokens progressively decreases in deeper layers during action generation. Motivated by this observation, we propose \textbf{DeepVision-VLA}, built on a \textbf{Vision-Language Mixture-of-Transformers (VL-MoT)} framework. This framework enables shared attention between the vision foundation model and the VLA backbone, injecting multi-level visual features from the vision expert into deeper layers of the VLA backbone to enhance visual representations for precise and complex manipulation. In addition, we introduce \textbf{Action-Guided Visual Pruning (AGVP)}, which leverages shallow-layer attention to prune irrelevant visual tokens while preserving task-relevant ones, reinforcing critical visual cues for manipulation with minimal computational overhead. DeepVision-VLA outperforms prior state-of-the-art methods by 9.0\% and 7.5\% on simulated and real-world tasks, respectively, providing new insights for the design of visually enhanced VLA models.
title Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.15618