Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiao, Zheng, Xiang, Gao, Yifeng, Xia, Xinyu, Wang, Yixu, Wang, Xin, Sun, Ye, Zhao, Yunhan, Wen, Ming, Li, Jiayu, Chen, Zixing, Gong, Xun, Liu, Yi, Li, Yige, Wu, Yutao, Wang, Cong, Sun, Jun, Cao, Yixin, Chen, Zhineng, Chen, Jingjing, Gui, Tao, Zhang, Qi, Wu, Zuxuan, Qiu, Xipeng, Huang, Xuanjing, Zhang, Tiehua, Wei, Zhipeng, Wang, Kun, Li, Xinfeng, Huang, Hanxun, Erfani, Sarah, Bailey, James, Wang, Jianping, Xiao, Chaowei, He, Ran, Li, Bo, Ma, Xingjun, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
by: Cui, Kaiyuan, et al.
Published: (2026)
by: Cui, Kaiyuan, et al.
Published: (2026)
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
Detecting Backdoor Samples in Contrastive Language Image Pretraining
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
by: Li, Juncheng, et al.
Published: (2025)
by: Li, Juncheng, et al.
Published: (2025)
Internal Safety Collapse in Frontier Large Language Models
by: Wu, Yutao, et al.
Published: (2026)
by: Wu, Yutao, et al.
Published: (2026)
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
by: Zheng, Xiang, et al.
Published: (2026)
by: Zheng, Xiang, et al.
Published: (2026)
HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models
by: Chen, Zixing, et al.
Published: (2026)
by: Chen, Zixing, et al.
Published: (2026)
DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning
by: Gao, Yifeng, et al.
Published: (2025)
by: Gao, Yifeng, et al.
Published: (2025)
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
by: Li, Yige, et al.
Published: (2026)
by: Li, Yige, et al.
Published: (2026)
Expose Before You Defend: Unifying and Enhancing Backdoor Defenses via Exposed Models
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
AudioMosaic: Contrastive Masked Audio Representation Learning
by: Huang, Hanxun, et al.
Published: (2026)
by: Huang, Hanxun, et al.
Published: (2026)
Downstream Transfer Attack: Adversarial Attacks on Downstream Models with Pre-trained Vision Transformers
by: Zheng, Weijie, et al.
Published: (2024)
by: Zheng, Weijie, et al.
Published: (2024)
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
by: Ma, Xingjun, et al.
Published: (2025)
by: Ma, Xingjun, et al.
Published: (2025)
End-to-End Anti-Backdoor Learning on Images and Time Series
by: Jiang, Yujing, et al.
Published: (2024)
by: Jiang, Yujing, et al.
Published: (2024)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
by: Chen, Yunhao, et al.
Published: (2025)
by: Chen, Yunhao, et al.
Published: (2025)
A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
by: Ma, Xingjun, et al.
Published: (2026)
by: Ma, Xingjun, et al.
Published: (2026)
FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds
by: Wang, Qizhou, et al.
Published: (2025)
by: Wang, Qizhou, et al.
Published: (2025)
ModelLock: Locking Your Model With a Spell
by: Gao, Yifeng, et al.
Published: (2024)
by: Gao, Yifeng, et al.
Published: (2024)
T2UE: Generating Unlearnable Examples from Text Descriptions
by: Ma, Xingjun, et al.
Published: (2025)
by: Ma, Xingjun, et al.
Published: (2025)
Collision‐free flocking control of uncertain multi‐agent systems: Combining current learning adaptive control and projection operator
by: Ximing Wang, et al.
Published: (2025)
by: Ximing Wang, et al.
Published: (2025)
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
by: Wang, Weiyi, et al.
Published: (2026)
by: Wang, Weiyi, et al.
Published: (2026)
StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
Towards Context-Invariant Safety Alignment for Large Language Models
by: Wang, Yixu, et al.
Published: (2026)
by: Wang, Yixu, et al.
Published: (2026)
SentGuard: Sentence-Level Streaming Guardrails for Large Language Models
by: Yu, Jiaqi, et al.
Published: (2026)
by: Yu, Jiaqi, et al.
Published: (2026)
Geometry-Guided Adversarial Prompt Detection via Curvature and Local Intrinsic Dimension
by: Yung, Canaan, et al.
Published: (2025)
by: Yung, Canaan, et al.
Published: (2025)
Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs
by: Hasanebrahimi, Afsaneh, et al.
Published: (2026)
by: Hasanebrahimi, Afsaneh, et al.
Published: (2026)
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning
by: Sun, Bowen, et al.
Published: (2026)
by: Sun, Bowen, et al.
Published: (2026)
AdvQDet: Detecting Query-Based Adversarial Attacks with Adversarial Contrastive Prompt Tuning
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
by: Li, Jiayu, et al.
Published: (2025)
by: Li, Jiayu, et al.
Published: (2025)
Consistency Purification: Effective and Efficient Diffusion Purification towards Certified Robustness
by: Li, Yiquan, et al.
Published: (2024)
by: Li, Yiquan, et al.
Published: (2024)
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
by: Zhao, Yunhan, et al.
Published: (2024)
by: Zhao, Yunhan, et al.
Published: (2024)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
by: Du, Mengfei, et al.
Published: (2024)
by: Du, Mengfei, et al.
Published: (2024)
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
by: Wu, Yutao, et al.
Published: (2025)
by: Wu, Yutao, et al.
Published: (2025)
Similar Items
-
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
by: Li, Yige, et al.
Published: (2024) -
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
by: Cui, Kaiyuan, et al.
Published: (2026) -
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
by: Huang, Hanxun, et al.
Published: (2025) -
Detecting Backdoor Samples in Contrastive Language Image Pretraining
by: Huang, Hanxun, et al.
Published: (2025) -
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
by: Li, Juncheng, et al.
Published: (2025)