Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Taowen, Han, Cheng, Liang, James Chenhao, Yang, Wenhao, Liu, Dongfang, Zhang, Luna Xinyu, Wang, Qifan, Luo, Jiebo, Tang, Ruixiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908474224410624
author Wang, Taowen
Han, Cheng
Liang, James Chenhao
Yang, Wenhao
Liu, Dongfang
Zhang, Luna Xinyu
Wang, Qifan
Luo, Jiebo
Tang, Ruixiang
author_facet Wang, Taowen
Han, Cheng
Liang, James Chenhao
Yang, Wenhao
Liu, Dongfang
Zhang, Luna Xinyu
Wang, Qifan
Luo, Jiebo
Tang, Ruixiang
contents Recently in robotics, Vision-Language-Action (VLA) models have emerged as a transformative approach, enabling robots to execute complex tasks by integrating visual and linguistic inputs within an end-to-end learning framework. Despite their significant capabilities, VLA models introduce new attack surfaces. This paper systematically evaluates their robustness. Recognizing the unique demands of robotic execution, our attack objectives target the inherent spatial and functional characteristics of robotic systems. In particular, we introduce two untargeted attack objectives that leverage spatial foundations to destabilize robotic actions, and a targeted attack objective that manipulates the robotic trajectory. Additionally, we design an adversarial patch generation approach that places a small, colorful patch within the camera's view, effectively executing the attack in both digital and physical environments. Our evaluation reveals a marked degradation in task success rates, with up to a 100\% reduction across a suite of simulated robotic tasks, highlighting critical security gaps in current VLA architectures. By unveiling these vulnerabilities and proposing actionable evaluation metrics, we advance both the understanding and enhancement of safety for VLA-based robotic systems, underscoring the necessity for continuously developing robust defense strategies prior to physical-world deployments.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13587
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics
Wang, Taowen
Han, Cheng
Liang, James Chenhao
Yang, Wenhao
Liu, Dongfang
Zhang, Luna Xinyu
Wang, Qifan
Luo, Jiebo
Tang, Ruixiang
Robotics
Artificial Intelligence
Recently in robotics, Vision-Language-Action (VLA) models have emerged as a transformative approach, enabling robots to execute complex tasks by integrating visual and linguistic inputs within an end-to-end learning framework. Despite their significant capabilities, VLA models introduce new attack surfaces. This paper systematically evaluates their robustness. Recognizing the unique demands of robotic execution, our attack objectives target the inherent spatial and functional characteristics of robotic systems. In particular, we introduce two untargeted attack objectives that leverage spatial foundations to destabilize robotic actions, and a targeted attack objective that manipulates the robotic trajectory. Additionally, we design an adversarial patch generation approach that places a small, colorful patch within the camera's view, effectively executing the attack in both digital and physical environments. Our evaluation reveals a marked degradation in task success rates, with up to a 100\% reduction across a suite of simulated robotic tasks, highlighting critical security gaps in current VLA architectures. By unveiling these vulnerabilities and proposing actionable evaluation metrics, we advance both the understanding and enhancement of safety for VLA-based robotic systems, underscoring the necessity for continuously developing robust defense strategies prior to physical-world deployments.
title Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2411.13587