RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Chan Hee, Blukis, Valts, Tremblay, Jonathan, Tyree, Stephen, Su, Yu, Birchfield, Stan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
by: Singh, Ishika, et al.
Published: (2025)
by: Singh, Ishika, et al.
Published: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)
by: Yuan, Wentao, et al.
Published: (2024)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
Neural Implicit Representation for Building Digital Twins of Unknown Articulated Objects
by: Weng, Yijia, et al.
Published: (2024)
by: Weng, Yijia, et al.
Published: (2024)
GRS: Generating Robotic Simulation Tasks from Real-World Images
by: Zook, Alex, et al.
Published: (2024)
by: Zook, Alex, et al.
Published: (2024)
3D-MVP: 3D Multiview Pretraining for Robotic Manipulation
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
by: Ma, Xianzheng, et al.
Published: (2026)
by: Ma, Xianzheng, et al.
Published: (2026)
Snap-it, Tap-it, Splat-it: Tactile-Informed 3D Gaussian Splatting for Reconstructing Challenging Surfaces
by: Comi, Mauro, et al.
Published: (2024)
by: Comi, Mauro, et al.
Published: (2024)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
cuRoboV2: Dynamics-Aware Motion Generation with Depth-Fused Distance Fields for High-DoF Robots
by: Sundaralingam, Balakumar, et al.
Published: (2026)
by: Sundaralingam, Balakumar, et al.
Published: (2026)
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
by: Wang, Siyin, et al.
Published: (2025)
by: Wang, Siyin, et al.
Published: (2025)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
by: Liu, Fanfan, et al.
Published: (2024)
by: Liu, Fanfan, et al.
Published: (2024)
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
by: Mu, Yao, et al.
Published: (2025)
by: Mu, Yao, et al.
Published: (2025)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
by: Goetting, Dylan, et al.
Published: (2024)
by: Goetting, Dylan, et al.
Published: (2024)
FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
by: Wen, Bowen, et al.
Published: (2023)
by: Wen, Bowen, et al.
Published: (2023)
Spatial RoboGrasp: Generalized Robotic Grasping Control Policy
by: Huang, Yiqi, et al.
Published: (2025)
by: Huang, Yiqi, et al.
Published: (2025)
Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
by: Wen, Bowen, et al.
Published: (2025)
by: Wen, Bowen, et al.
Published: (2025)
3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds
by: Sun, Fan-Yun, et al.
Published: (2025)
by: Sun, Fan-Yun, et al.
Published: (2025)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
by: Li, Huiqiong, et al.
Published: (2026)
by: Li, Huiqiong, et al.
Published: (2026)
RoboPlayground: Democratizing Robotic Evaluation through Structured Physical Domains
by: Wang, Yi Ru, et al.
Published: (2026)
by: Wang, Yi Ru, et al.
Published: (2026)
Robot Policy Evaluation for Sim-to-Real Transfer: A Benchmarking Perspective
by: Yang, Xuning, et al.
Published: (2025)
by: Yang, Xuning, et al.
Published: (2025)
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version)
by: Mu, Yao, et al.
Published: (2024)
by: Mu, Yao, et al.
Published: (2024)
Precise Robot Command Understanding Using Grammar-Constrained Large Language Models
by: Huo, Xinyun, et al.
Published: (2026)
by: Huo, Xinyun, et al.
Published: (2026)
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
by: Yang, Xuning, et al.
Published: (2026)
by: Yang, Xuning, et al.
Published: (2026)
RVT-2: Learning Precise Manipulation from Few Demonstrations
by: Goyal, Ankit, et al.
Published: (2024)
by: Goyal, Ankit, et al.
Published: (2024)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
by: Jia, Baoxiong, et al.
Published: (2024)
by: Jia, Baoxiong, et al.
Published: (2024)
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
by: Ma, Boyi, et al.
Published: (2025)
by: Ma, Boyi, et al.
Published: (2025)
Vision-Language Interpreter for Robot Task Planning
by: Shirai, Keisuke, et al.
Published: (2023)
by: Shirai, Keisuke, et al.
Published: (2023)
PRISM: Preference Refinement via Implicit Scene Modeling for 3D Vision-Language Preference-Based Reinforcement Learning
by: Sun, Yirong, et al.
Published: (2025)
by: Sun, Yirong, et al.
Published: (2025)
Towards Multimodal Social Conversations with Robots: Using Vision-Language Models
by: Janssens, Ruben, et al.
Published: (2025)
by: Janssens, Ruben, et al.
Published: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
by: Xia, Zhongyu, et al.
Published: (2026)
by: Xia, Zhongyu, et al.
Published: (2026)
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
by: Min, Cheolhong, et al.
Published: (2026)
by: Min, Cheolhong, et al.
Published: (2026)
Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations
by: Sitdhipol, Supawich, et al.
Published: (2025)
by: Sitdhipol, Supawich, et al.
Published: (2025)
Stable Language Guidance for Vision-Language-Action Models
by: Zhan, Zhihao, et al.
Published: (2026)
by: Zhan, Zhihao, et al.
Published: (2026)
3D-VLA: A 3D Vision-Language-Action Generative World Model
by: Zhen, Haoyu, et al.
Published: (2024)
by: Zhen, Haoyu, et al.
Published: (2024)
Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding
by: Kabir, Imran, et al.
Published: (2025)
by: Kabir, Imran, et al.
Published: (2025)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
by: Kranti, Chalamalasetti, et al.
Published: (2026)
by: Kranti, Chalamalasetti, et al.
Published: (2026)
Similar Items
-
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
by: Singh, Ishika, et al.
Published: (2025) -
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024) -
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2025) -
Neural Implicit Representation for Building Digital Twins of Unknown Articulated Objects
by: Weng, Yijia, et al.
Published: (2024) -
GRS: Generating Robotic Simulation Tasks from Real-World Images
by: Zook, Alex, et al.
Published: (2024)