Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Hao, Xiao, Erjia, Wang, Yichi, Yu, Chengyuan, Sun, Mengshu, Zhang, Qiang, Cao, Jiahang, Guo, Yijie, Liu, Ning, Xu, Kaidi, Zhang, Jize, Shen, Chao, Torr, Philip, Gu, Jindong, Xu, Renjing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models
by: Cheng, Hao, et al.
Published: (2025)
by: Cheng, Hao, et al.
Published: (2025)
Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
by: Cheng, Hao, et al.
Published: (2025)
by: Cheng, Hao, et al.
Published: (2025)
Gaining the Sparse Rewards by Exploring Lottery Tickets in Spiking Neural Network
by: Cheng, Hao, et al.
Published: (2023)
by: Cheng, Hao, et al.
Published: (2023)
Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
Prompting Multi-Modal Tokens to Enhance End-to-End Autonomous Driving Imitation Learning with LLMs
by: Duan, Yiqun, et al.
Published: (2024)
by: Duan, Yiqun, et al.
Published: (2024)
SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents
by: Cheng, Hao, et al.
Published: (2026)
by: Cheng, Hao, et al.
Published: (2026)
Event Masked Autoencoder: Point-wise Action Recognition with Event-Based Cameras
by: Sun, Jingkai, et al.
Published: (2025)
by: Sun, Jingkai, et al.
Published: (2025)
An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
by: Luo, Haochen, et al.
Published: (2024)
by: Luo, Haochen, et al.
Published: (2024)
AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control
by: Xu, Peng, et al.
Published: (2026)
by: Xu, Peng, et al.
Published: (2026)
LenslessFace: An End-to-End Optimized Lensless System for Privacy-Preserving Face Verification
by: Cai, Xin, et al.
Published: (2024)
by: Cai, Xin, et al.
Published: (2024)
Multi-Floor Zero-Shot Object Navigation Policy
by: Zhang, Lingfeng, et al.
Published: (2024)
by: Zhang, Lingfeng, et al.
Published: (2024)
DNCASR: End-to-End Training for Speaker-Attributed ASR
by: Zheng, Xianrui, et al.
Published: (2025)
by: Zheng, Xianrui, et al.
Published: (2025)
ES-Parkour: Advanced Robot Parkour with Bio-inspired Event Camera and Spiking Neural Network
by: Zhang, Qiang, et al.
Published: (2025)
by: Zhang, Qiang, et al.
Published: (2025)
TriHelper: Zero-Shot Object Navigation with Dynamic Assistance
by: Zhang, Lingfeng, et al.
Published: (2024)
by: Zhang, Lingfeng, et al.
Published: (2024)
Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion
by: Zhang, Jiaru, et al.
Published: (2026)
by: Zhang, Jiaru, et al.
Published: (2026)
CryptoFace: End-to-End Encrypted Face Recognition
by: Ao, Wei, et al.
Published: (2025)
by: Ao, Wei, et al.
Published: (2025)
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
Speaker Adaptation for Quantised End-to-End ASR Models
by: Zhao, Qiuming, et al.
Published: (2024)
by: Zhao, Qiuming, et al.
Published: (2024)
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
by: Yu, Yonghui, et al.
Published: (2025)
by: Yu, Yonghui, et al.
Published: (2025)
Fully Spiking Neural Network for Legged Robots
by: Jiang, Xiaoyang, et al.
Published: (2023)
by: Jiang, Xiaoyang, et al.
Published: (2023)
Modality-Composable Diffusion Policy via Inference-Time Distribution-level Composition
by: Cao, Jiahang, et al.
Published: (2025)
by: Cao, Jiahang, et al.
Published: (2025)
Influencer Backdoor Attack on Semantic Segmentation
by: Lan, Haoheng, et al.
Published: (2023)
by: Lan, Haoheng, et al.
Published: (2023)
Can an Individual Manipulate the Collective Decisions of Multi-Agents?
by: Liu, Fengyuan, et al.
Published: (2025)
by: Liu, Fengyuan, et al.
Published: (2025)
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
by: Wu, Yixuan, et al.
Published: (2025)
by: Wu, Yixuan, et al.
Published: (2025)
Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Distillation-PPO: A Novel Two-Stage Reinforcement Learning Framework for Humanoid Robot Perceptive Locomotion
by: Zhang, Qiang, et al.
Published: (2025)
by: Zhang, Qiang, et al.
Published: (2025)
Multi-Agent End-to-End Vulnerability Management for Mitigating Recurring Vulnerabilities
by: Zheng, Zelong, et al.
Published: (2026)
by: Zheng, Zelong, et al.
Published: (2026)
Distributionally Robust Control with End-to-End Statistically Guaranteed Metric Learning
by: Wu, Jingyi, et al.
Published: (2025)
by: Wu, Jingyi, et al.
Published: (2025)
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
by: Zhang, Kaidi, et al.
Published: (2026)
by: Zhang, Kaidi, et al.
Published: (2026)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
DSPO: An End-to-End Framework for Direct Sorted Portfolio Construction
by: Zhong, Jianyuan, et al.
Published: (2024)
by: Zhong, Jianyuan, et al.
Published: (2024)
Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models
by: Xing, Songlong, et al.
Published: (2026)
by: Xing, Songlong, et al.
Published: (2026)
RRAM-Based Bio-Inspired Circuits for Mobile Epileptic Correlation Extraction and Seizure Prediction
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
by: Rezaei, Razieh, et al.
Published: (2024)
by: Rezaei, Razieh, et al.
Published: (2024)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
by: Wang, Linbo, et al.
Published: (2026)
by: Wang, Linbo, et al.
Published: (2026)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
by: Zhao, Qiuming, et al.
Published: (2024)
by: Zhao, Qiuming, et al.
Published: (2024)
MoSA: Motion-constrained Stress Adaptation for Mitigating Real-to-Sim Gap in Continuum Dynamics via Learning Residual Anisotropy
by: Wang, Jiaxu, et al.
Published: (2026)
by: Wang, Jiaxu, et al.
Published: (2026)
LiPS: Large-Scale Humanoid Robot Reinforcement Learning with Parallel-Series Structures
by: Zhang, Qiang, et al.
Published: (2025)
by: Zhang, Qiang, et al.
Published: (2025)
Similar Items
-
Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
by: Cheng, Hao, et al.
Published: (2024) -
Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models
by: Cheng, Hao, et al.
Published: (2025) -
Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model
by: Cheng, Hao, et al.
Published: (2024) -
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
by: Cheng, Hao, et al.
Published: (2025) -
Gaining the Sparse Rewards by Exploring Lottery Tickets in Spiking Neural Network
by: Cheng, Hao, et al.
Published: (2023)