ESPIRE: A Diagnostic Benchmark for Embodied Spatial Reasoning of Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Yanpeng, Ding, Wentao, Li, Hongtao, Jia, Baoxiong, Zheng, Zilong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
di: Yang, Yandan, et al.
Pubblicazione: (2024)
di: Yang, Yandan, et al.
Pubblicazione: (2024)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent
di: Yang, Yandan, et al.
Pubblicazione: (2025)
di: Yang, Yandan, et al.
Pubblicazione: (2025)
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
di: Zhang, Jiyao, et al.
Pubblicazione: (2026)
di: Zhang, Jiyao, et al.
Pubblicazione: (2026)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
di: Xing, Shuo, et al.
Pubblicazione: (2024)
di: Xing, Shuo, et al.
Pubblicazione: (2024)
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
di: Guruprasad, Pranav, et al.
Pubblicazione: (2024)
di: Guruprasad, Pranav, et al.
Pubblicazione: (2024)
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
di: Liu, Yu, et al.
Pubblicazione: (2024)
di: Liu, Yu, et al.
Pubblicazione: (2024)
ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark
di: Dang, Ronghao, et al.
Pubblicazione: (2025)
di: Dang, Ronghao, et al.
Pubblicazione: (2025)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
di: Jia, Baoxiong, et al.
Pubblicazione: (2024)
di: Jia, Baoxiong, et al.
Pubblicazione: (2024)
Octopus: Embodied Vision-Language Programmer from Environmental Feedback
di: Yang, Jingkang, et al.
Pubblicazione: (2023)
di: Yang, Jingkang, et al.
Pubblicazione: (2023)
HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents
di: Tomilin, Tristan, et al.
Pubblicazione: (2025)
di: Tomilin, Tristan, et al.
Pubblicazione: (2025)
ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
di: Chahe, Amirhosein, et al.
Pubblicazione: (2025)
di: Chahe, Amirhosein, et al.
Pubblicazione: (2025)
GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
di: Lu, Guanxing, et al.
Pubblicazione: (2025)
di: Lu, Guanxing, et al.
Pubblicazione: (2025)
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
di: Zhang, Junjie, et al.
Pubblicazione: (2024)
di: Zhang, Junjie, et al.
Pubblicazione: (2024)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
di: Zhang, Zhengshen, et al.
Pubblicazione: (2025)
di: Zhang, Zhengshen, et al.
Pubblicazione: (2025)
GrabS: Generative Embodied Agent for 3D Object Segmentation without Scene Supervision
di: Zhang, Zihui, et al.
Pubblicazione: (2025)
di: Zhang, Zihui, et al.
Pubblicazione: (2025)
ArtGS: Building Interactable Replicas of Complex Articulated Objects via Gaussian Splatting
di: Liu, Yu, et al.
Pubblicazione: (2025)
di: Liu, Yu, et al.
Pubblicazione: (2025)
Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied Agents
di: Zhang, Zhizhen, et al.
Pubblicazione: (2025)
di: Zhang, Zhizhen, et al.
Pubblicazione: (2025)
Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models
di: Lee, Seungjae, et al.
Pubblicazione: (2025)
di: Lee, Seungjae, et al.
Pubblicazione: (2025)
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
di: NVIDIA, et al.
Pubblicazione: (2025)
di: NVIDIA, et al.
Pubblicazione: (2025)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
di: Zhao, Qingqing, et al.
Pubblicazione: (2025)
di: Zhao, Qingqing, et al.
Pubblicazione: (2025)
Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V
di: Zhi, Peiyuan, et al.
Pubblicazione: (2024)
di: Zhi, Peiyuan, et al.
Pubblicazione: (2024)
DistortBench: Benchmarking Vision Language Models on Image Distortion Identification
di: Goyal, Divyanshu, et al.
Pubblicazione: (2026)
di: Goyal, Divyanshu, et al.
Pubblicazione: (2026)
Tactile Modality Fusion for Vision-Language-Action Models
di: Morissette, Charlotte, et al.
Pubblicazione: (2026)
di: Morissette, Charlotte, et al.
Pubblicazione: (2026)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
di: Li, Chengmeng, et al.
Pubblicazione: (2025)
di: Li, Chengmeng, et al.
Pubblicazione: (2025)
Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning
di: Ganai, Milan, et al.
Pubblicazione: (2026)
di: Ganai, Milan, et al.
Pubblicazione: (2026)
Learning Human-Humanoid Coordination for Collaborative Object Carrying
di: Du, Yushi, et al.
Pubblicazione: (2025)
di: Du, Yushi, et al.
Pubblicazione: (2025)
AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
di: AgiBot-World-Contributors, et al.
Pubblicazione: (2025)
di: AgiBot-World-Contributors, et al.
Pubblicazione: (2025)
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
di: Wu, Tao, et al.
Pubblicazione: (2025)
di: Wu, Tao, et al.
Pubblicazione: (2025)
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
di: Cheang, Chi-Lam, et al.
Pubblicazione: (2024)
di: Cheang, Chi-Lam, et al.
Pubblicazione: (2024)
PVI: Plug-in Visual Injection for Vision-Language-Action Models
di: Zhang, Zezhou, et al.
Pubblicazione: (2026)
di: Zhang, Zezhou, et al.
Pubblicazione: (2026)
Improving Vision-Language-Action Model with Online Reinforcement Learning
di: Guo, Yanjiang, et al.
Pubblicazione: (2025)
di: Guo, Yanjiang, et al.
Pubblicazione: (2025)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
di: Ling, Yiran, et al.
Pubblicazione: (2026)
di: Ling, Yiran, et al.
Pubblicazione: (2026)
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
di: Zhu, Haoyi, et al.
Pubblicazione: (2024)
di: Zhu, Haoyi, et al.
Pubblicazione: (2024)
Understanding Particles From Video: Property Estimation of Granular Materials via Visuo-Haptic Learning
di: Zhang, Zeqing, et al.
Pubblicazione: (2024)
di: Zhang, Zeqing, et al.
Pubblicazione: (2024)
Test-Time Training for Visual Foresight Vision-Language-Action Models
di: Park, Sangwu, et al.
Pubblicazione: (2026)
di: Park, Sangwu, et al.
Pubblicazione: (2026)
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
di: Huang, Siyuan, et al.
Pubblicazione: (2025)
di: Huang, Siyuan, et al.
Pubblicazione: (2025)
Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning
di: Ding, Runyu, et al.
Pubblicazione: (2024)
di: Ding, Runyu, et al.
Pubblicazione: (2024)
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
di: Wu, Tao, et al.
Pubblicazione: (2024)
di: Wu, Tao, et al.
Pubblicazione: (2024)
AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
di: Vasa, Santosh, et al.
Pubblicazione: (2025)
di: Vasa, Santosh, et al.
Pubblicazione: (2025)
Documenti analoghi
-
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
di: Yang, Yandan, et al.
Pubblicazione: (2024) -
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
di: Chen, Boyuan, et al.
Pubblicazione: (2024) -
SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent
di: Yang, Yandan, et al.
Pubblicazione: (2025) -
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
di: Zhang, Jiyao, et al.
Pubblicazione: (2026) -
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
di: Xing, Shuo, et al.
Pubblicazione: (2024)