State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Ji, Jiang, Wenbo, Lin, Yansong, Liu, Yijing, Zhang, Ruichen, Lu, Guomin, Chen, Aiguo, Han, Xinshuo, Li, Hongwei, Niyato, Dusit
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911359997837312
author Guo, Ji
Jiang, Wenbo
Lin, Yansong
Liu, Yijing
Zhang, Ruichen
Lu, Guomin
Chen, Aiguo
Han, Xinshuo
Li, Hongwei
Niyato, Dusit
author_facet Guo, Ji
Jiang, Wenbo
Lin, Yansong
Liu, Yijing
Zhang, Ruichen
Lu, Guomin
Chen, Aiguo
Han, Xinshuo
Li, Hongwei
Niyato, Dusit
contents Vision-Language-Action (VLA) models are widely deployed in safety-critical embodied AI applications such as robotics. However, their complex multimodal interactions also expose new security vulnerabilities. In this paper, we investigate a backdoor threat in VLA models, where malicious inputs cause targeted misbehavior while preserving performance on clean data. Existing backdoor methods predominantly rely on inserting visible triggers into visual modality, which suffer from poor robustness and low insusceptibility in real-world settings due to environmental variability. To overcome these limitations, we introduce the State Backdoor, a novel and practical backdoor attack that leverages the robot arm's initial state as the trigger. To optimize trigger for insusceptibility and effectiveness, we design a Preference-guided Genetic Algorithm (PGA) that efficiently searches the state space for minimal yet potent triggers. Extensive experiments on five representative VLA models and five real-world tasks show that our method achieves over 90% attack success rate without affecting benign task performance, revealing an underexplored vulnerability in embodied AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04266
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space
Guo, Ji
Jiang, Wenbo
Lin, Yansong
Liu, Yijing
Zhang, Ruichen
Lu, Guomin
Chen, Aiguo
Han, Xinshuo
Li, Hongwei
Niyato, Dusit
Cryptography and Security
Machine Learning
Vision-Language-Action (VLA) models are widely deployed in safety-critical embodied AI applications such as robotics. However, their complex multimodal interactions also expose new security vulnerabilities. In this paper, we investigate a backdoor threat in VLA models, where malicious inputs cause targeted misbehavior while preserving performance on clean data. Existing backdoor methods predominantly rely on inserting visible triggers into visual modality, which suffer from poor robustness and low insusceptibility in real-world settings due to environmental variability. To overcome these limitations, we introduce the State Backdoor, a novel and practical backdoor attack that leverages the robot arm's initial state as the trigger. To optimize trigger for insusceptibility and effectiveness, we design a Preference-guided Genetic Algorithm (PGA) that efficiently searches the state space for minimal yet potent triggers. Extensive experiments on five representative VLA models and five real-world tasks show that our method achieves over 90% attack success rate without affecting benign task performance, revealing an underexplored vulnerability in embodied AI systems.
title State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2601.04266