SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Boyuan, Xu, Zhuo, Kirmani, Sean, Ichter, Brian, Driess, Danny, Florence, Pete, Sadigh, Dorsa, Guibas, Leonidas, Xia, Fei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
Chain of Code: Reasoning with a Language Model-Augmented Code Emulator
by: Li, Chengshu, et al.
Published: (2023)
by: Li, Chengshu, et al.
Published: (2023)
Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
by: Bai, Qianqian, et al.
Published: (2025)
by: Bai, Qianqian, et al.
Published: (2025)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024)
by: Bharadhwaj, Homanga, et al.
Published: (2024)
Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization
by: Yang, Jonathan, et al.
Published: (2025)
by: Yang, Jonathan, et al.
Published: (2025)
Vision Language Models are In-Context Value Learners
by: Ma, Yecheng Jason, et al.
Published: (2024)
by: Ma, Yecheng Jason, et al.
Published: (2024)
Unified Video Action Model
by: Li, Shuang, et al.
Published: (2025)
by: Li, Shuang, et al.
Published: (2025)
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
by: Kranti, Chalamalasetti, et al.
Published: (2026)
by: Kranti, Chalamalasetti, et al.
Published: (2026)
ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2025)
by: Grannen, Jennifer, et al.
Published: (2025)
MotIF: Motion Instruction Fine-tuning
by: Hwang, Minyoung, et al.
Published: (2024)
by: Hwang, Minyoung, et al.
Published: (2024)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
Vocal Sandbox: Continual Learning and Adaptation for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2024)
by: Grannen, Jennifer, et al.
Published: (2024)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
by: Ahn, Michael, et al.
Published: (2024)
by: Ahn, Michael, et al.
Published: (2024)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
by: Goetting, Dylan, et al.
Published: (2024)
by: Goetting, Dylan, et al.
Published: (2024)
ALOHA Unleashed: A Simple Recipe for Robot Dexterity
by: Zhao, Tony Z., et al.
Published: (2024)
by: Zhao, Tony Z., et al.
Published: (2024)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning
by: Englmeier, Stefan, et al.
Published: (2026)
by: Englmeier, Stefan, et al.
Published: (2026)
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
Action-Free Reasoning for Policy Generalization
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation
by: Liu, Jiaxing, et al.
Published: (2026)
by: Liu, Jiaxing, et al.
Published: (2026)
What's the Move? Hybrid Imitation Learning via Salient Points
by: Sundaresan, Priya, et al.
Published: (2024)
by: Sundaresan, Priya, et al.
Published: (2024)
SAGE: Bridging Semantic and Actionable Parts for GEneralizable Manipulation of Articulated Objects
by: Geng, Haoran, et al.
Published: (2023)
by: Geng, Haoran, et al.
Published: (2023)
HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies
by: Xie, Amber, et al.
Published: (2026)
by: Xie, Amber, et al.
Published: (2026)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
by: Xia, Zhongyu, et al.
Published: (2026)
by: Xia, Zhongyu, et al.
Published: (2026)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
ESPIRE: A Diagnostic Benchmark for Embodied Spatial Reasoning of Vision-Language Models
by: Zhao, Yanpeng, et al.
Published: (2026)
by: Zhao, Yanpeng, et al.
Published: (2026)
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
by: Ye, Angen, et al.
Published: (2025)
by: Ye, Angen, et al.
Published: (2025)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
by: Fang, Jiading
Published: (2025)
by: Fang, Jiading
Published: (2025)
SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation
by: Wang, Qianxu, et al.
Published: (2023)
by: Wang, Qianxu, et al.
Published: (2023)
Rodrigues Network for Learning Robot Actions
by: Zhang, Jialiang, et al.
Published: (2025)
by: Zhang, Jialiang, et al.
Published: (2025)
RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation
by: Qin, Zheng, et al.
Published: (2025)
by: Qin, Zheng, et al.
Published: (2025)
DOT-Sim: Differentiable Optical Tactile Simulation with Precise Real-to-Sim Physical Calibration
by: You, Yang, et al.
Published: (2026)
by: You, Yang, et al.
Published: (2026)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
How to Train Your Robots? The Impact of Demonstration Modality on Imitation Learning
by: Li, Haozhuo, et al.
Published: (2025)
by: Li, Haozhuo, et al.
Published: (2025)
Similar Items
-
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024) -
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023) -
Chain of Code: Reasoning with a Language Model-Augmented Code Emulator
by: Li, Chengshu, et al.
Published: (2023) -
Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
by: Bai, Qianqian, et al.
Published: (2025) -
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
by: Hong, Yining, et al.
Published: (2026)