AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
Fuente:
arXiv
Guardado en:
| Autores principales: | Ying, Zonghao, Wang, Le, Xiao, Yisong, Wang, Jiakai, Ma, Yuqing, Guo, Jinyang, Yin, Zhenfei, Zhang, Mingchuan, Liu, Aishan, Liu, Xianglong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
por: Ying, Zonghao, et al.
Publicado: (2024)
por: Ying, Zonghao, et al.
Publicado: (2024)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
por: Ying, Zonghao, et al.
Publicado: (2024)
por: Ying, Zonghao, et al.
Publicado: (2024)
Evolving Deception: When Agents Evolve, Deception Wins
por: Ying, Zonghao, et al.
Publicado: (2026)
por: Ying, Zonghao, et al.
Publicado: (2026)
Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw
por: Ying, Zonghao, et al.
Publicado: (2026)
por: Ying, Zonghao, et al.
Publicado: (2026)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
por: Ying, Zonghao, et al.
Publicado: (2026)
por: Ying, Zonghao, et al.
Publicado: (2026)
RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic
por: Wang, Le, et al.
Publicado: (2025)
por: Wang, Le, et al.
Publicado: (2025)
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents
por: Yeke, Doguhuan, et al.
Publicado: (2026)
por: Yeke, Doguhuan, et al.
Publicado: (2026)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
por: Yin, Sheng, et al.
Publicado: (2024)
por: Yin, Sheng, et al.
Publicado: (2024)
Compromising Embodied Agents with Contextual Backdoor Attacks
por: Liu, Aishan, et al.
Publicado: (2024)
por: Liu, Aishan, et al.
Publicado: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
por: Ying, Zonghao, et al.
Publicado: (2024)
por: Ying, Zonghao, et al.
Publicado: (2024)
Beyond Model Jailbreak: Systematic Dissection of the "Ten DeadlySins" in Embodied Intelligence
por: Huang, Yuhang, et al.
Publicado: (2025)
por: Huang, Yuhang, et al.
Publicado: (2025)
The Shawshank Redemption of Embodied AI: Understanding and Benchmarking Indirect Environmental Jailbreaks
por: Li, Chunyang, et al.
Publicado: (2025)
por: Li, Chunyang, et al.
Publicado: (2025)
Achieving the Safety and Security of the End-to-End AV Pipeline
por: Curran, Noah T., et al.
Publicado: (2024)
por: Curran, Noah T., et al.
Publicado: (2024)
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
por: Li, Xiao, et al.
Publicado: (2026)
por: Li, Xiao, et al.
Publicado: (2026)
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
Property-Guided Cyber-Physical Reduction and Surrogation for Safety Analysis in Robotic Vehicles
por: Sayom, Nazmus Shakib, et al.
Publicado: (2025)
por: Sayom, Nazmus Shakib, et al.
Publicado: (2025)
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
por: Zou, Quanchen, et al.
Publicado: (2026)
por: Zou, Quanchen, et al.
Publicado: (2026)
Drones that Think on their Feet: Sudden Landing Decisions with Embodied AI
por: Barbosa, Diego Ortiz, et al.
Publicado: (2025)
por: Barbosa, Diego Ortiz, et al.
Publicado: (2025)
Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks
por: Xing, Wenpeng, et al.
Publicado: (2025)
por: Xing, Wenpeng, et al.
Publicado: (2025)
DLP: towards active defense against backdoor attacks with decoupled learning process
por: Ying, Zonghao, et al.
Publicado: (2024)
por: Ying, Zonghao, et al.
Publicado: (2024)
NBA: defensive distillation for backdoor removal via neural behavior alignment
por: Ying, Zonghao, et al.
Publicado: (2024)
por: Ying, Zonghao, et al.
Publicado: (2024)
RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors
por: Bai, Fengshuo, et al.
Publicado: (2024)
por: Bai, Fengshuo, et al.
Publicado: (2024)
Propagating Unsafe Actions in LLM Controlled Multi-Robot Collaboration via Single Robot Compromise
por: Huang, Zhen, et al.
Publicado: (2026)
por: Huang, Zhen, et al.
Publicado: (2026)
Revisiting Adversarial Perception Attacks and Defense Methods on Autonomous Driving Systems
por: Chen, Cheng, et al.
Publicado: (2025)
por: Chen, Cheng, et al.
Publicado: (2025)
PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification
por: Li, Hongwei, et al.
Publicado: (2025)
por: Li, Hongwei, et al.
Publicado: (2025)
AdvGrasp: Adversarial Attacks on Robotic Grasping from a Physical Perspective
por: Wang, Xiaofei, et al.
Publicado: (2025)
por: Wang, Xiaofei, et al.
Publicado: (2025)
Challenges in the Safety-Security Co-Assurance of Collaborative Industrial Robots
por: Gleirscher, Mario, et al.
Publicado: (2020)
por: Gleirscher, Mario, et al.
Publicado: (2020)
CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity
por: Yu, Zhengmin, et al.
Publicado: (2024)
por: Yu, Zhengmin, et al.
Publicado: (2024)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
por: Liu, Zhe, et al.
Publicado: (2026)
por: Liu, Zhe, et al.
Publicado: (2026)
Bayesian Methods for Trust in Collaborative Multi-Agent Autonomy
por: Hallyburton, R. Spencer, et al.
Publicado: (2024)
por: Hallyburton, R. Spencer, et al.
Publicado: (2024)
ARACNE: An LLM-Based Autonomous Shell Pentesting Agent
por: Nieponice, Tomas, et al.
Publicado: (2025)
por: Nieponice, Tomas, et al.
Publicado: (2025)
Not What You Asked For: Typographic Attacks in Household Robot Manipulation
por: Iranmanesh, Ali, et al.
Publicado: (2026)
por: Iranmanesh, Ali, et al.
Publicado: (2026)
Procedimiento de auditoría de ciberseguridad para sistemas autónomos: metodología, amenazas y mitigaciones
por: Campazas-Vega, Adrián, et al.
Publicado: (2025)
por: Campazas-Vega, Adrián, et al.
Publicado: (2025)
Channel State Information Analysis for Jamming Attack Detection in Static and Dynamic UAV Networks -- An Experimental Study
por: Mykytyn, Pavlo, et al.
Publicado: (2025)
por: Mykytyn, Pavlo, et al.
Publicado: (2025)
SoK: Cybersecurity Assessment of Humanoid Ecosystem
por: Surve, Priyanka Prakash, et al.
Publicado: (2025)
por: Surve, Priyanka Prakash, et al.
Publicado: (2025)
Uncertainty-Aware 3D Position Refinement for Multi-UAV Systems
por: Alamleh, Hosam, et al.
Publicado: (2026)
por: Alamleh, Hosam, et al.
Publicado: (2026)
Offensive Robot Cybersecurity
por: Mayoral-Vilches, Víctor
Publicado: (2025)
por: Mayoral-Vilches, Víctor
Publicado: (2025)
Ejemplares similares
-
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
por: Ying, Zonghao, et al.
Publicado: (2025) -
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
por: Ying, Zonghao, et al.
Publicado: (2024) -
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
por: Ying, Zonghao, et al.
Publicado: (2024) -
Evolving Deception: When Agents Evolve, Deception Wins
por: Ying, Zonghao, et al.
Publicado: (2026) -
Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw
por: Ying, Zonghao, et al.
Publicado: (2026)