Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Xuancun, Chen, Jiaxiang, Xiao, Shilin, Jin, Zizhi, Chen, Zhangrui, Yu, Hanwen, Qian, Bohan, Zhou, Ruochen, Ji, Xiaoyu, Xu, Wenyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909970529779712
author Lu, Xuancun
Chen, Jiaxiang
Xiao, Shilin
Jin, Zizhi
Chen, Zhangrui
Yu, Hanwen
Qian, Bohan
Zhou, Ruochen
Ji, Xiaoyu
Xu, Wenyuan
author_facet Lu, Xuancun
Chen, Jiaxiang
Xiao, Shilin
Jin, Zizhi
Chen, Zhangrui
Yu, Hanwen
Qian, Bohan
Zhou, Ruochen
Ji, Xiaoyu
Xu, Wenyuan
contents Vision-Language-Action (VLA) models revolutionize robotic systems by enabling end-to-end perception-to-action pipelines that integrate multiple sensory modalities, such as visual signals processed by cameras and auditory signals captured by microphones. This multi-modality integration allows VLA models to interpret complex, real-world environments using diverse sensor data streams. Given the fact that VLA-based systems heavily rely on the sensory input, the security of VLA models against physical-world sensor attacks remains critically underexplored. To address this gap, we present the first systematic study of physical sensor attacks against VLAs, quantifying the influence of sensor attacks and investigating the defenses for VLA models. We introduce a novel "Real-Sim-Real" framework that automatically simulates physics-based sensor attack vectors, including six attacks targeting cameras and two targeting microphones, and validates them on real robotic systems. Through large-scale evaluations across various VLA architectures and tasks under varying attack parameters, we demonstrate significant vulnerabilities, with susceptibility patterns that reveal critical dependencies on task types and model designs. We further develop an adversarial-training-based defense that enhances VLA robustness against out-of-distribution physical perturbations caused by sensor attacks while preserving model performance. Our findings expose an urgent need for standardized robustness benchmarks and mitigation strategies to secure VLA deployments in safety-critical environments.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10008
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks
Lu, Xuancun
Chen, Jiaxiang
Xiao, Shilin
Jin, Zizhi
Chen, Zhangrui
Yu, Hanwen
Qian, Bohan
Zhou, Ruochen
Ji, Xiaoyu
Xu, Wenyuan
Robotics
Artificial Intelligence
Vision-Language-Action (VLA) models revolutionize robotic systems by enabling end-to-end perception-to-action pipelines that integrate multiple sensory modalities, such as visual signals processed by cameras and auditory signals captured by microphones. This multi-modality integration allows VLA models to interpret complex, real-world environments using diverse sensor data streams. Given the fact that VLA-based systems heavily rely on the sensory input, the security of VLA models against physical-world sensor attacks remains critically underexplored. To address this gap, we present the first systematic study of physical sensor attacks against VLAs, quantifying the influence of sensor attacks and investigating the defenses for VLA models. We introduce a novel "Real-Sim-Real" framework that automatically simulates physics-based sensor attack vectors, including six attacks targeting cameras and two targeting microphones, and validates them on real robotic systems. Through large-scale evaluations across various VLA architectures and tasks under varying attack parameters, we demonstrate significant vulnerabilities, with susceptibility patterns that reveal critical dependencies on task types and model designs. We further develop an adversarial-training-based defense that enhances VLA robustness against out-of-distribution physical perturbations caused by sensor attacks while preserving model performance. Our findings expose an urgent need for standardized robustness benchmarks and mitigation strategies to secure VLA deployments in safety-critical environments.
title Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2511.10008