SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Zhan, Simon Sinong, Liu, Yao, Wang, Philip, Wang, Zinan, Wang, Qineng, Peng, Yiyan, Ruan, Zhian, Shi, Xiangyu, Cao, Xinyu, Yang, Frank, Wang, Kangrui, Shao, Huajie, Li, Manling, Zhu, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
by: Li, Manling, et al.
Published: (2024)
by: Li, Manling, et al.
Published: (2024)
NetBench: A Large-Scale and Comprehensive Network Traffic Benchmark Dataset for Foundation Models
by: Qian, Chen, et al.
Published: (2024)
by: Qian, Chen, et al.
Published: (2024)
Inverse Delayed Reinforcement Learning
by: Zhan, Simon Sinong, et al.
Published: (2024)
by: Zhan, Simon Sinong, et al.
Published: (2024)
Case Study: Runtime Safety Verification of Neural Network Controlled System
by: Yang, Frank, et al.
Published: (2024)
by: Yang, Frank, et al.
Published: (2024)
Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
by: Zhan, Simon Sinong, et al.
Published: (2025)
by: Zhan, Simon Sinong, et al.
Published: (2025)
Lens: A Knowledge-Guided Foundation Model for Network Traffic
by: Li, Xiaochang, et al.
Published: (2024)
by: Li, Xiaochang, et al.
Published: (2024)
From Twitter to Reasoner: Understand Mobility Travel Modes and Sentiment Using Large Language Models
by: Ruan, Kangrui, et al.
Published: (2024)
by: Ruan, Kangrui, et al.
Published: (2024)
ODESteer: A Unified ODE-Based Steering Framework for LLM Alignment
by: Zhao, Hongjue, et al.
Published: (2026)
by: Zhao, Hongjue, et al.
Published: (2026)
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
by: Wang, Kangrui, et al.
Published: (2025)
by: Wang, Kangrui, et al.
Published: (2025)
ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning
by: Chen, Hanyang, et al.
Published: (2025)
by: Chen, Hanyang, et al.
Published: (2025)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL
by: Ni, Zhenyang, et al.
Published: (2026)
by: Ni, Zhenyang, et al.
Published: (2026)
Empowering Autonomous Driving with Large Language Models: A Safety Perspective
by: Wang, Yixuan, et al.
Published: (2023)
by: Wang, Yixuan, et al.
Published: (2023)
Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?
by: Wang, Qineng, et al.
Published: (2024)
by: Wang, Qineng, et al.
Published: (2024)
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping
by: Zhan, Simon Sinong, et al.
Published: (2024)
by: Zhan, Simon Sinong, et al.
Published: (2024)
MindCube: Spatial Mental Modeling from Limited Views
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
by: Ju, Ruofei, et al.
Published: (2026)
by: Ju, Ruofei, et al.
Published: (2026)
Strategic Optimization and Challenges of Large Language Models in Object-Oriented Programming
by: Wang, Zinan
Published: (2024)
by: Wang, Zinan
Published: (2024)
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents
by: X, Tencent Robotics, et al.
Published: (2026)
by: X, Tencent Robotics, et al.
Published: (2026)
Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
by: Ye, Ruimeng, et al.
Published: (2025)
by: Ye, Ruimeng, et al.
Published: (2025)
Training Cross-Morphology Embodied AI Agents: From Practical Challenges to Theoretical Foundations
by: Liu, Shaoshan, et al.
Published: (2025)
by: Liu, Shaoshan, et al.
Published: (2025)
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
by: Zhang, Pingyue, et al.
Published: (2026)
by: Zhang, Pingyue, et al.
Published: (2026)
Distributed Invariant Unscented Kalman Filter based on Inverse Covariance Intersection with Intermittent Measurements
by: Ruan, Zhian, et al.
Published: (2024)
by: Ruan, Zhian, et al.
Published: (2024)
RAGEN-2: Reasoning Collapse in Agentic RL
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
Safety of Embodied Navigation: A Survey
by: Wang, Zixia, et al.
Published: (2025)
by: Wang, Zixia, et al.
Published: (2025)
Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation
by: Wang, Ning, et al.
Published: (2025)
by: Wang, Ning, et al.
Published: (2025)
SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
by: Chen, Ruolin, et al.
Published: (2025)
by: Chen, Ruolin, et al.
Published: (2025)
See and Think: Embodied Agent in Virtual Environment
by: Zhao, Zhonghan, et al.
Published: (2023)
by: Zhao, Zhonghan, et al.
Published: (2023)
SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
Your Language Model Secretly Contains Personality Subnetworks
by: Ye, Ruimeng, et al.
Published: (2026)
by: Ye, Ruimeng, et al.
Published: (2026)
Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression
by: Peng, Jingyu, et al.
Published: (2025)
by: Peng, Jingyu, et al.
Published: (2025)
Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own
by: Ye, Weirui, et al.
Published: (2023)
by: Ye, Weirui, et al.
Published: (2023)
FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making
by: Wang, Yucen, et al.
Published: (2025)
by: Wang, Yucen, et al.
Published: (2025)
Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Universal Actions for Enhanced Embodied Foundation Models
by: Zheng, Jinliang, et al.
Published: (2025)
by: Zheng, Jinliang, et al.
Published: (2025)
Differentially Private Synthetic Data via APIs 3: Using Simulators Instead of Foundation Model
by: Lin, Zinan, et al.
Published: (2025)
by: Lin, Zinan, et al.
Published: (2025)
RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration
by: Tan, Huajie, et al.
Published: (2025)
by: Tan, Huajie, et al.
Published: (2025)
Similar Items
-
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
by: Yang, Rui, et al.
Published: (2025) -
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
by: Li, Manling, et al.
Published: (2024) -
NetBench: A Large-Scale and Comprehensive Network Traffic Benchmark Dataset for Foundation Models
by: Qian, Chen, et al.
Published: (2024) -
Inverse Delayed Reinforcement Learning
by: Zhan, Simon Sinong, et al.
Published: (2024) -
Case Study: Runtime Safety Verification of Neural Network Controlled System
by: Yang, Frank, et al.
Published: (2024)