EARBench: Towards Evaluating Physical Risk Awareness for Task Planning of Foundation Model-based Embodied AI Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Zihao, Wu, Bingzhe, Zhang, Zhengyou, Han, Lei, Liu, Qingshan, Wu, Baoyuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VDC: Versatile Data Cleanser based on Visual-Linguistic Inconsistency by Multimodal Large Language Models
di: Zhu, Zihao, et al.
Pubblicazione: (2023)
di: Zhu, Zihao, et al.
Pubblicazione: (2023)
MFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning
di: Zhang, Min, et al.
Pubblicazione: (2024)
di: Zhang, Min, et al.
Pubblicazione: (2024)
The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents
di: Wu, Baoyuan, et al.
Pubblicazione: (2026)
di: Wu, Baoyuan, et al.
Pubblicazione: (2026)
Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model
di: Ma, Huan, et al.
Pubblicazione: (2024)
di: Ma, Huan, et al.
Pubblicazione: (2024)
ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models
di: Lai, Zhenglin, et al.
Pubblicazione: (2026)
di: Lai, Zhenglin, et al.
Pubblicazione: (2026)
MADRA: Multi-Agent Debate for Risk-Aware Embodied Planning
di: Wang, Junjian, et al.
Pubblicazione: (2025)
di: Wang, Junjian, et al.
Pubblicazione: (2025)
HMGIE: Hierarchical and Multi-Grained Inconsistency Evaluation for Vision-Language Data Cleansing
di: Zhu, Zihao, et al.
Pubblicazione: (2024)
di: Zhu, Zihao, et al.
Pubblicazione: (2024)
ET-Plan-Bench: Embodied Task-level Planning Benchmark Towards Spatial-Temporal Cognition with Foundation Models
di: Zhang, Lingfeng, et al.
Pubblicazione: (2024)
di: Zhang, Lingfeng, et al.
Pubblicazione: (2024)
Unveiling Covert Toxicity in Multimodal Data via Toxicity Association Graphs: A Graph-Based Metric and Interpretable Detection Framework
di: Wu, Guanzong, et al.
Pubblicazione: (2026)
di: Wu, Guanzong, et al.
Pubblicazione: (2026)
BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation
di: Zhu, Zihao, et al.
Pubblicazione: (2026)
di: Zhu, Zihao, et al.
Pubblicazione: (2026)
STMA: A Spatio-Temporal Memory Agent for Long-Horizon Embodied Task Planning
di: Lei, Mingcong, et al.
Pubblicazione: (2025)
di: Lei, Mingcong, et al.
Pubblicazione: (2025)
AgenticCache: Cache-Driven Asynchronous Planning for Embodied AI Agents
di: Kim, Hojoon, et al.
Pubblicazione: (2026)
di: Kim, Hojoon, et al.
Pubblicazione: (2026)
A Survey on Robotics with Foundation Models: toward Embodied AI
di: Xu, Zhiyuan, et al.
Pubblicazione: (2024)
di: Xu, Zhiyuan, et al.
Pubblicazione: (2024)
SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents
di: Zhan, Simon Sinong, et al.
Pubblicazione: (2025)
di: Zhan, Simon Sinong, et al.
Pubblicazione: (2025)
The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents
di: Wang, Ziyu, et al.
Pubblicazione: (2026)
di: Wang, Ziyu, et al.
Pubblicazione: (2026)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
di: Zhu, Zihao, et al.
Pubblicazione: (2025)
di: Zhu, Zihao, et al.
Pubblicazione: (2025)
BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning
di: Ma, Xiaoyu, et al.
Pubblicazione: (2026)
di: Ma, Xiaoyu, et al.
Pubblicazione: (2026)
Physical Reasoning and Object Planning for Household Embodied Agents
di: Agrawal, Ayush, et al.
Pubblicazione: (2023)
di: Agrawal, Ayush, et al.
Pubblicazione: (2023)
ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement Learning
di: Zhou, Weijie, et al.
Pubblicazione: (2025)
di: Zhou, Weijie, et al.
Pubblicazione: (2025)
Automatic Cognitive Task Generation for In-Situ Evaluation of Embodied Agents
di: He, Xinyi, et al.
Pubblicazione: (2026)
di: He, Xinyi, et al.
Pubblicazione: (2026)
Plan Verification for LLM-Based Embodied Task Completion Agents
di: Hariharan, Ananth, et al.
Pubblicazione: (2025)
di: Hariharan, Ananth, et al.
Pubblicazione: (2025)
MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents
di: Jiang, Dongming, et al.
Pubblicazione: (2026)
di: Jiang, Dongming, et al.
Pubblicazione: (2026)
SDA-PLANNER: State-Dependency Aware Adaptive Planner for Embodied Task Planning
di: Shen, Zichao, et al.
Pubblicazione: (2025)
di: Shen, Zichao, et al.
Pubblicazione: (2025)
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
di: Ju, Ruofei, et al.
Pubblicazione: (2026)
di: Ju, Ruofei, et al.
Pubblicazione: (2026)
Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents
di: Kim, Byeonghwi, et al.
Pubblicazione: (2023)
di: Kim, Byeonghwi, et al.
Pubblicazione: (2023)
Attacks in Adversarial Machine Learning: A Systematic Survey from the Life-cycle Perspective
di: Wu, Baoyuan, et al.
Pubblicazione: (2023)
di: Wu, Baoyuan, et al.
Pubblicazione: (2023)
To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models
di: Zhu, Zihao, et al.
Pubblicazione: (2025)
di: Zhu, Zihao, et al.
Pubblicazione: (2025)
SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents
di: Zhu, Kunlun, et al.
Pubblicazione: (2025)
di: Zhu, Kunlun, et al.
Pubblicazione: (2025)
Embodied AI Agents: Modeling the World
di: Fung, Pascale, et al.
Pubblicazione: (2025)
di: Fung, Pascale, et al.
Pubblicazione: (2025)
Towards Objectively Benchmarking Social Intelligence for Language Agents at Action Level
di: Wang, Chenxu, et al.
Pubblicazione: (2024)
di: Wang, Chenxu, et al.
Pubblicazione: (2024)
Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
di: Wang, Zihao, et al.
Pubblicazione: (2023)
di: Wang, Zihao, et al.
Pubblicazione: (2023)
Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems
di: Liu, Bin
Pubblicazione: (2025)
di: Liu, Bin
Pubblicazione: (2025)
TPS-Bench: Evaluating AI Agents' Tool Planning \& Scheduling Abilities in Compounding Tasks
di: Xu, Hanwen, et al.
Pubblicazione: (2025)
di: Xu, Hanwen, et al.
Pubblicazione: (2025)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
di: Yin, Sheng, et al.
Pubblicazione: (2024)
di: Yin, Sheng, et al.
Pubblicazione: (2024)
Towards Responsible Generative AI: A Reference Architecture for Designing Foundation Model based Agents
di: Lu, Qinghua, et al.
Pubblicazione: (2023)
di: Lu, Qinghua, et al.
Pubblicazione: (2023)
Training Cross-Morphology Embodied AI Agents: From Practical Challenges to Theoretical Foundations
di: Liu, Shaoshan, et al.
Pubblicazione: (2025)
di: Liu, Shaoshan, et al.
Pubblicazione: (2025)
SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation
di: Chen, Ziyi, et al.
Pubblicazione: (2025)
di: Chen, Ziyi, et al.
Pubblicazione: (2025)
Reliable Poisoned Sample Detection against Backdoor Attacks Enhanced by Sharpness Aware Minimization
di: Zhang, Mingda, et al.
Pubblicazione: (2024)
di: Zhang, Mingda, et al.
Pubblicazione: (2024)
TANGO: Training-free Embodied AI Agents for Open-world Tasks
di: Ziliotto, Filippo, et al.
Pubblicazione: (2024)
di: Ziliotto, Filippo, et al.
Pubblicazione: (2024)
LLM3:Large Language Model-based Task and Motion Planning with Motion Failure Reasoning
di: Wang, Shu, et al.
Pubblicazione: (2024)
di: Wang, Shu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
VDC: Versatile Data Cleanser based on Visual-Linguistic Inconsistency by Multimodal Large Language Models
di: Zhu, Zihao, et al.
Pubblicazione: (2023) -
MFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning
di: Zhang, Min, et al.
Pubblicazione: (2024) -
The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents
di: Wu, Baoyuan, et al.
Pubblicazione: (2026) -
Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model
di: Ma, Huan, et al.
Pubblicazione: (2024) -
ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models
di: Lai, Zhenglin, et al.
Pubblicazione: (2026)