Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Ning, Yan, Zihan, Li, Weiyang, Ma, Chuan, Chen, He, Xiang, Tao |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
par: Chen, Ruolin, et autres
Publié: (2025)
par: Chen, Ruolin, et autres
Publié: (2025)
BEDI: A Comprehensive Benchmark for Evaluating Embodied Agents on UAVs
par: Guo, Mingning, et autres
Publié: (2025)
par: Guo, Mingning, et autres
Publié: (2025)
A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents
par: Huang, Yuting, et autres
Publié: (2025)
par: Huang, Yuting, et autres
Publié: (2025)
EmboCoach-Bench: Benchmarking AI Agents on Developing Embodied Robots
par: Lei, Zixing, et autres
Publié: (2026)
par: Lei, Zixing, et autres
Publié: (2026)
EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment
par: Gao, Chen, et autres
Publié: (2024)
par: Gao, Chen, et autres
Publié: (2024)
EmbodiedGovBench: A Benchmark for Governance, Recovery, and Upgrade Safety in Embodied Agent Systems
par: Qin, Xue, et autres
Publié: (2026)
par: Qin, Xue, et autres
Publié: (2026)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
par: Li, Manling, et autres
Publié: (2024)
par: Li, Manling, et autres
Publié: (2024)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
par: Yin, Sheng, et autres
Publié: (2024)
par: Yin, Sheng, et autres
Publié: (2024)
ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning
par: Chen, Hanyang, et autres
Publié: (2025)
par: Chen, Hanyang, et autres
Publié: (2025)
PULSE: Privileged Knowledge Transfer from Rich to Deployable Sensors for Embodied Multi-Sensory Learning
par: Zhao, Zihan, et autres
Publié: (2025)
par: Zhao, Zihan, et autres
Publié: (2025)
AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
par: Luo, Hanjun, et autres
Publié: (2025)
par: Luo, Hanjun, et autres
Publié: (2025)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
par: Evtimov, Ivan, et autres
Publié: (2025)
par: Evtimov, Ivan, et autres
Publié: (2025)
SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents
par: Zhan, Simon Sinong, et autres
Publié: (2025)
par: Zhan, Simon Sinong, et autres
Publié: (2025)
CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space
par: Zhao, Yong, et autres
Publié: (2025)
par: Zhao, Yong, et autres
Publié: (2025)
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
par: Ju, Ruofei, et autres
Publié: (2026)
par: Ju, Ruofei, et autres
Publié: (2026)
Safety of Embodied Navigation: A Survey
par: Wang, Zixia, et autres
Publié: (2025)
par: Wang, Zixia, et autres
Publié: (2025)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
par: Yuan, Tongxin, et autres
Publié: (2024)
par: Yuan, Tongxin, et autres
Publié: (2024)
Safety Control of Service Robots with LLMs and Embodied Knowledge Graphs
par: Qi, Yong, et autres
Publié: (2024)
par: Qi, Yong, et autres
Publié: (2024)
EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents
par: Juneja, Gurusha, et autres
Publié: (2026)
par: Juneja, Gurusha, et autres
Publié: (2026)
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
par: Hong, Zhepei, et autres
Publié: (2026)
par: Hong, Zhepei, et autres
Publié: (2026)
See and Think: Embodied Agent in Virtual Environment
par: Zhao, Zhonghan, et autres
Publié: (2023)
par: Zhao, Zhonghan, et autres
Publié: (2023)
LoTa-Bench: Benchmarking Language-oriented Task Planners for Embodied Agents
par: Choi, Jae-Woo, et autres
Publié: (2024)
par: Choi, Jae-Woo, et autres
Publié: (2024)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
par: Yang, Rui, et autres
Publié: (2025)
par: Yang, Rui, et autres
Publié: (2025)
RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic
par: Wang, Le, et autres
Publié: (2025)
par: Wang, Le, et autres
Publié: (2025)
Multi-agent Embodied AI: Advances and Future Directions
par: Feng, Zhaohan, et autres
Publié: (2025)
par: Feng, Zhaohan, et autres
Publié: (2025)
Security of AI Agents
par: He, Yifeng, et autres
Publié: (2024)
par: He, Yifeng, et autres
Publié: (2024)
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
par: Guo, Weiyang, et autres
Publié: (2025)
par: Guo, Weiyang, et autres
Publié: (2025)
FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
par: Zeng, Zhiyuan, et autres
Publié: (2025)
par: Zeng, Zhiyuan, et autres
Publié: (2025)
SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World
par: Zhang, Jiaqi, et autres
Publié: (2024)
par: Zhang, Jiaqi, et autres
Publié: (2024)
SkillTester: Benchmarking Utility and Security of Agent Skills
par: Wang, Leye, et autres
Publié: (2026)
par: Wang, Leye, et autres
Publié: (2026)
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
par: Wang, Siyin, et autres
Publié: (2024)
par: Wang, Siyin, et autres
Publié: (2024)
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
par: Li, Xiao, et autres
Publié: (2026)
par: Li, Xiao, et autres
Publié: (2026)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
par: Li, Miles Q., et autres
Publié: (2026)
par: Li, Miles Q., et autres
Publié: (2026)
Embodied CoT Distillation From LLM To Off-the-shelf Agents
par: Choi, Wonje, et autres
Publié: (2024)
par: Choi, Wonje, et autres
Publié: (2024)
SAGE: A Service Agent Graph-guided Evaluation Benchmark
par: Shi, Ling, et autres
Publié: (2026)
par: Shi, Ling, et autres
Publié: (2026)
Safety2Drive: Safety-Critical Scenario Benchmark for the Evaluation of Autonomous Driving
par: Li, Jingzheng, et autres
Publié: (2025)
par: Li, Jingzheng, et autres
Publié: (2025)
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
par: Yang, Zhi, et autres
Publié: (2026)
par: Yang, Zhi, et autres
Publié: (2026)
The Safety Challenge of World Models for Embodied AI Agents: A Review
par: Baraldi, Lorenzo, et autres
Publié: (2025)
par: Baraldi, Lorenzo, et autres
Publié: (2025)
Embodied AI Agents: Modeling the World
par: Fung, Pascale, et autres
Publié: (2025)
par: Fung, Pascale, et autres
Publié: (2025)
Human-centered In-building Embodied Delivery Benchmark
par: Xu, Zhuoqun, et autres
Publié: (2024)
par: Xu, Zhuoqun, et autres
Publié: (2024)
Documents similaires
-
SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
par: Chen, Ruolin, et autres
Publié: (2025) -
BEDI: A Comprehensive Benchmark for Evaluating Embodied Agents on UAVs
par: Guo, Mingning, et autres
Publié: (2025) -
A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents
par: Huang, Yuting, et autres
Publié: (2025) -
EmboCoach-Bench: Benchmarking AI Agents on Developing Embodied Robots
par: Lei, Zixing, et autres
Publié: (2026) -
EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment
par: Gao, Chen, et autres
Publié: (2024)