Selective Perception for Robot: Task-Aware Attention in Multimodal VLA

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Son, Young-Chae, Lee, Jung-Woo, Choi, Yoon-Ji, Ko, Dae-Kwan, Lim, Soo-Chul
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915801782550528
author Son, Young-Chae
Lee, Jung-Woo
Choi, Yoon-Ji
Ko, Dae-Kwan
Lim, Soo-Chul
author_facet Son, Young-Chae
Lee, Jung-Woo
Choi, Yoon-Ji
Ko, Dae-Kwan
Lim, Soo-Chul
contents In robotics, Vision-Language-Action (VLA) models that integrate diverse multimodal signals from multi-view inputs have emerged as an effective approach. However, most prior work adopts static fusion that processes all visual inputs uniformly, which incurs unnecessary computational overhead and allows task-irrelevant background information to act as noise. Inspired by the principles of human active perception, we propose a dynamic information fusion framework designed to maximize the efficiency and robustness of VLA models. Our approach introduces a lightweight adaptive routing architecture that analyzes the current text prompt and observations from a wrist-mounted camera in real-time to predict the task-relevance of multiple camera views. By conditionally attenuating computations for views with low informational utility and selectively providing only essential visual features to the policy network, Our framework achieves computation efficiency proportional to task relevance. Furthermore, to efficiently secure large-scale annotation data for router training, we established an automated labeling pipeline utilizing Vision-Language Models (VLMs) to minimize data collection and annotation costs. Experimental results in real-world robotic manipulation scenarios demonstrate that the proposed approach achieves significant improvements in both inference efficiency and control performance compared to existing VLA models, validating the effectiveness and practicality of dynamic information fusion in resource-constrained, real-time robot control environments.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15543
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Selective Perception for Robot: Task-Aware Attention in Multimodal VLA
Son, Young-Chae
Lee, Jung-Woo
Choi, Yoon-Ji
Ko, Dae-Kwan
Lim, Soo-Chul
Robotics
In robotics, Vision-Language-Action (VLA) models that integrate diverse multimodal signals from multi-view inputs have emerged as an effective approach. However, most prior work adopts static fusion that processes all visual inputs uniformly, which incurs unnecessary computational overhead and allows task-irrelevant background information to act as noise. Inspired by the principles of human active perception, we propose a dynamic information fusion framework designed to maximize the efficiency and robustness of VLA models. Our approach introduces a lightweight adaptive routing architecture that analyzes the current text prompt and observations from a wrist-mounted camera in real-time to predict the task-relevance of multiple camera views. By conditionally attenuating computations for views with low informational utility and selectively providing only essential visual features to the policy network, Our framework achieves computation efficiency proportional to task relevance. Furthermore, to efficiently secure large-scale annotation data for router training, we established an automated labeling pipeline utilizing Vision-Language Models (VLMs) to minimize data collection and annotation costs. Experimental results in real-world robotic manipulation scenarios demonstrate that the proposed approach achieves significant improvements in both inference efficiency and control performance compared to existing VLA models, validating the effectiveness and practicality of dynamic information fusion in resource-constrained, real-time robot control environments.
title Selective Perception for Robot: Task-Aware Attention in Multimodal VLA
topic Robotics
url https://arxiv.org/abs/2602.15543