ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Yining, Liu, Jiageng, Yin, Han, Li, Manling, Guibas, Leonidas, Fei-Fei, Li, Wu, Jiajun, Choi, Yejin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
Planning with the Views via Scene Self-Exploration
by: Wang, Kangrui, et al.
Published: (2026)
by: Wang, Kangrui, et al.
Published: (2026)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
by: Li, Manling, et al.
Published: (2024)
by: Li, Manling, et al.
Published: (2024)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
by: Kuang, Yuxuan, et al.
Published: (2024)
by: Kuang, Yuxuan, et al.
Published: (2024)
Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2024)
by: Jia, Xiaosong, et al.
Published: (2024)
Closed Loop Interactive Embodied Reasoning for Robot Manipulation
by: Nazarczuk, Michal, et al.
Published: (2024)
by: Nazarczuk, Michal, et al.
Published: (2024)
Rodrigues Network for Learning Robot Actions
by: Zhang, Jialiang, et al.
Published: (2025)
by: Zhang, Jialiang, et al.
Published: (2025)
StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception
by: Han, Evans, et al.
Published: (2026)
by: Han, Evans, et al.
Published: (2026)
Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
by: Hong, Yining, et al.
Published: (2025)
by: Hong, Yining, et al.
Published: (2025)
PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning
by: Li, Haoyang, et al.
Published: (2026)
by: Li, Haoyang, et al.
Published: (2026)
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
by: Zhang, Jiyao, et al.
Published: (2026)
by: Zhang, Jiyao, et al.
Published: (2026)
Vision-Language Navigation with Embodied Intelligence: A Survey
by: Gao, Peng, et al.
Published: (2024)
by: Gao, Peng, et al.
Published: (2024)
World Action Models: The Next Frontier in Embodied AI
by: Wang, Siyin, et al.
Published: (2026)
by: Wang, Siyin, et al.
Published: (2026)
Open-Loop Planning, Closed-Loop Verification: Speculative Verification for VLA
by: Wang, Zihua, et al.
Published: (2026)
by: Wang, Zihua, et al.
Published: (2026)
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
by: Lang, Xiaolei, et al.
Published: (2026)
by: Lang, Xiaolei, et al.
Published: (2026)
EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence
by: Wang, Xinjie, et al.
Published: (2025)
by: Wang, Xinjie, et al.
Published: (2025)
Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy
by: Wu, Pengyuan, et al.
Published: (2026)
by: Wu, Pengyuan, et al.
Published: (2026)
Intelligent Spatial Perception by Building Hierarchical 3D Scene Graphs for Indoor Scenarios with the Help of LLMs
by: Cheng, Yao, et al.
Published: (2025)
by: Cheng, Yao, et al.
Published: (2025)
Towards Efficient LLM Grounding for Embodied Multi-Agent Collaboration
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Dejavu: Towards Experience Feedback Learning for Embodied Intelligence
by: Wu, Shaokai, et al.
Published: (2025)
by: Wu, Shaokai, et al.
Published: (2025)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
by: Sun, Qi, et al.
Published: (2024)
by: Sun, Qi, et al.
Published: (2024)
Segment-to-Act: Label-Noise-Robust Action-Prompted Video Segmentation Towards Embodied Intelligence
by: Li, Wenxin, et al.
Published: (2025)
by: Li, Wenxin, et al.
Published: (2025)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
by: Fang, Jiading
Published: (2025)
by: Fang, Jiading
Published: (2025)
Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
by: Yang, Jiashu, et al.
Published: (2025)
by: Yang, Jiashu, et al.
Published: (2025)
VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
SAGE: Bridging Semantic and Actionable Parts for GEneralizable Manipulation of Articulated Objects
by: Geng, Haoran, et al.
Published: (2023)
by: Geng, Haoran, et al.
Published: (2023)
SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
by: Renz, Katrin, et al.
Published: (2025)
by: Renz, Katrin, et al.
Published: (2025)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
by: Li, Zhenxin, et al.
Published: (2025)
by: Li, Zhenxin, et al.
Published: (2025)
MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
by: Hong, Yining, et al.
Published: (2024)
by: Hong, Yining, et al.
Published: (2024)
OnlineHOI: Towards Online Human-Object Interaction Generation and Perception
by: Ji, Yihong, et al.
Published: (2025)
by: Ji, Yihong, et al.
Published: (2025)
EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
by: Zou, Ding, et al.
Published: (2025)
by: Zou, Ding, et al.
Published: (2025)
Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop
by: Kerr, Justin, et al.
Published: (2025)
by: Kerr, Justin, et al.
Published: (2025)
PhysPart: Physically Plausible Part Completion for Interactable Objects
by: Luo, Rundong, et al.
Published: (2024)
by: Luo, Rundong, et al.
Published: (2024)
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
by: Dharmarajan, Karthik, et al.
Published: (2025)
by: Dharmarajan, Karthik, et al.
Published: (2025)
Bag-of-Word-Groups (BoWG): A Robust and Efficient Loop Closure Detection Method Under Perceptual Aliasing
by: Fei, Xiang, et al.
Published: (2025)
by: Fei, Xiang, et al.
Published: (2025)
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
by: Lei, Zixing, et al.
Published: (2026)
by: Lei, Zixing, et al.
Published: (2026)
Towards Dense and Accurate Radar Perception Via Efficient Cross-Modal Diffusion Model
by: Zhang, Ruibin, et al.
Published: (2024)
by: Zhang, Ruibin, et al.
Published: (2024)
Similar Items
-
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
by: Hong, Yining, et al.
Published: (2026) -
Planning with the Views via Scene Self-Exploration
by: Wang, Kangrui, et al.
Published: (2026) -
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
by: Wang, Qineng, et al.
Published: (2025) -
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
by: Li, Manling, et al.
Published: (2024) -
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)