The Shawshank Redemption of Embodied AI: Understanding and Benchmarking Indirect Environmental Jailbreaks
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Chunyang, Kang, Zifeng, Zhang, Junwei, Ma, Zhuo, Cheng, Anda, Li, Xinghua, Ma, Jianfeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CPA-RAG:Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models
di: Li, Chunyang, et al.
Pubblicazione: (2025)
di: Li, Chunyang, et al.
Pubblicazione: (2025)
Beyond Model Jailbreak: Systematic Dissection of the "Ten DeadlySins" in Embodied Intelligence
di: Huang, Yuhang, et al.
Pubblicazione: (2025)
di: Huang, Yuhang, et al.
Pubblicazione: (2025)
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents
di: Yeke, Doguhuan, et al.
Pubblicazione: (2026)
di: Yeke, Doguhuan, et al.
Pubblicazione: (2026)
Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks
di: Xing, Wenpeng, et al.
Pubblicazione: (2025)
di: Xing, Wenpeng, et al.
Pubblicazione: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
Drones that Think on their Feet: Sudden Landing Decisions with Embodied AI
di: Barbosa, Diego Ortiz, et al.
Pubblicazione: (2025)
di: Barbosa, Diego Ortiz, et al.
Pubblicazione: (2025)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
di: Yin, Sheng, et al.
Pubblicazione: (2024)
di: Yin, Sheng, et al.
Pubblicazione: (2024)
Benchmarking and Understanding Safety Risks in AI Character Platforms
di: Wei, Yiluo, et al.
Pubblicazione: (2025)
di: Wei, Yiluo, et al.
Pubblicazione: (2025)
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
di: Li, Xiao, et al.
Pubblicazione: (2026)
di: Li, Xiao, et al.
Pubblicazione: (2026)
Cumplimiento del Reglamento (UE) 2024/1689 en robótica y sistemas autónomos: una revisión sistemática de la literatura
di: Lorenzo, Yoana Pita
Pubblicazione: (2025)
di: Lorenzo, Yoana Pita
Pubblicazione: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
LightDefense: A Lightweight Uncertainty-Driven Defense against Jailbreaks via Shifted Token Distribution
di: Yang, Zhuoran, et al.
Pubblicazione: (2025)
di: Yang, Zhuoran, et al.
Pubblicazione: (2025)
EM-MIAs: Enhancing Membership Inference Attacks in Large Language Models through Ensemble Modeling
di: Song, Zichen, et al.
Pubblicazione: (2024)
di: Song, Zichen, et al.
Pubblicazione: (2024)
Adjustable AprilTags For Identity Secured Tasks
di: Li, Hao
Pubblicazione: (2025)
di: Li, Hao
Pubblicazione: (2025)
SoK: On the Semantic AI Security in Autonomous Driving
di: Shen, Junjie, et al.
Pubblicazione: (2022)
di: Shen, Junjie, et al.
Pubblicazione: (2022)
Understanding Users' Security and Privacy Concerns and Attitudes Towards Conversational AI Platforms
di: Ali, Mutahar, et al.
Pubblicazione: (2025)
di: Ali, Mutahar, et al.
Pubblicazione: (2025)
Global AI Governance Overview: Understanding Regulatory Requirements Across Global Jurisdictions
di: Kyrychenko, Mariia, et al.
Pubblicazione: (2025)
di: Kyrychenko, Mariia, et al.
Pubblicazione: (2025)
On the Feasibility of Fingerprinting Collaborative Robot Network Traffic
di: Tang, Cheng, et al.
Pubblicazione: (2023)
di: Tang, Cheng, et al.
Pubblicazione: (2023)
Diagnosis-guided Attack Recovery for Securing Robotic Vehicles from Sensor Deception Attacks
di: Dash, Pritam, et al.
Pubblicazione: (2022)
di: Dash, Pritam, et al.
Pubblicazione: (2022)
Revisiting Adversarial Perception Attacks and Defense Methods on Autonomous Driving Systems
di: Chen, Cheng, et al.
Pubblicazione: (2025)
di: Chen, Cheng, et al.
Pubblicazione: (2025)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
di: Xu, Zonghuan, et al.
Pubblicazione: (2025)
di: Xu, Zonghuan, et al.
Pubblicazione: (2025)
AdvGrasp: Adversarial Attacks on Robotic Grasping from a Physical Perspective
di: Wang, Xiaofei, et al.
Pubblicazione: (2025)
di: Wang, Xiaofei, et al.
Pubblicazione: (2025)
Achieving the Safety and Security of the End-to-End AV Pipeline
di: Curran, Noah T., et al.
Pubblicazione: (2024)
di: Curran, Noah T., et al.
Pubblicazione: (2024)
MM-AttacKG: A Multimodal Approach to Attack Graph Construction with Large Language Models
di: Zhang, Yongheng, et al.
Pubblicazione: (2025)
di: Zhang, Yongheng, et al.
Pubblicazione: (2025)
When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
di: Zhu, Changjia, et al.
Pubblicazione: (2025)
di: Zhu, Changjia, et al.
Pubblicazione: (2025)
ProvX: Generating Counterfactual-Driven Attack Explanations for Provenance-Based Detection
di: Wu, Weiheng, et al.
Pubblicazione: (2025)
di: Wu, Weiheng, et al.
Pubblicazione: (2025)
Blockchain in Environmental Sustainability Measures: a Survey
di: Vladucu, Maria-Victoria, et al.
Pubblicazione: (2024)
di: Vladucu, Maria-Victoria, et al.
Pubblicazione: (2024)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
di: Murphy, Brendan, et al.
Pubblicazione: (2025)
di: Murphy, Brendan, et al.
Pubblicazione: (2025)
TombRaider: Entering the Vault of History to Jailbreak Large Language Models
di: Ding, Junchen, et al.
Pubblicazione: (2025)
di: Ding, Junchen, et al.
Pubblicazione: (2025)
Formal Verification of Robustness and Resilience of Learning-Enabled State Estimation Systems
di: Huang, Wei, et al.
Pubblicazione: (2020)
di: Huang, Wei, et al.
Pubblicazione: (2020)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack
di: Zhuo, Terry Yue, et al.
Pubblicazione: (2026)
di: Zhuo, Terry Yue, et al.
Pubblicazione: (2026)
Careful About What App Promotion Ads Recommend! Detecting and Explaining Malware Promotion via App Promotion Graph
di: Ma, Shang, et al.
Pubblicazione: (2024)
di: Ma, Shang, et al.
Pubblicazione: (2024)
Understanding Cyber Threats Against the Universities, Colleges, and Schools
di: Lallie, Harjinder Singh, et al.
Pubblicazione: (2023)
di: Lallie, Harjinder Singh, et al.
Pubblicazione: (2023)
Red Team Redemption: A Structured Comparison of Open-Source Tools for Adversary Emulation
di: Landauer, Max, et al.
Pubblicazione: (2024)
di: Landauer, Max, et al.
Pubblicazione: (2024)
Understanding and Enhancing the Transferability of Jailbreaking Attacks
di: Lin, Runqi, et al.
Pubblicazione: (2025)
di: Lin, Runqi, et al.
Pubblicazione: (2025)
Infrastructure for Valuable, Tradable, and Verifiable Agent Memory
di: Li, Mengyuan, et al.
Pubblicazione: (2026)
di: Li, Mengyuan, et al.
Pubblicazione: (2026)
PerProb: Indirectly Evaluating Memorization in Large Language Models
di: Liao, Yihan, et al.
Pubblicazione: (2025)
di: Liao, Yihan, et al.
Pubblicazione: (2025)
Procedimiento de auditoría de ciberseguridad para sistemas autónomos: metodología, amenazas y mitigaciones
di: Campazas-Vega, Adrián, et al.
Pubblicazione: (2025)
di: Campazas-Vega, Adrián, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CPA-RAG:Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models
di: Li, Chunyang, et al.
Pubblicazione: (2025) -
Beyond Model Jailbreak: Systematic Dissection of the "Ten DeadlySins" in Embodied Intelligence
di: Huang, Yuhang, et al.
Pubblicazione: (2025) -
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
di: Ying, Zonghao, et al.
Pubblicazione: (2025) -
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents
di: Yeke, Doguhuan, et al.
Pubblicazione: (2026) -
Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks
di: Xing, Wenpeng, et al.
Pubblicazione: (2025)