SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Boyuan, Xu, Zhuo, Kirmani, Sean, Ichter, Brian, Driess, Danny, Florence, Pete, Sadigh, Dorsa, Guibas, Leonidas, Xia, Fei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909079534829568
author Chen, Boyuan
Xu, Zhuo
Kirmani, Sean
Ichter, Brian
Driess, Danny
Florence, Pete
Sadigh, Dorsa
Guibas, Leonidas
Xia, Fei
author_facet Chen, Boyuan
Xu, Zhuo
Kirmani, Sean
Ichter, Brian
Driess, Danny
Florence, Pete
Sadigh, Dorsa
Guibas, Leonidas
Xia, Fei
contents Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demonstrated remarkable performance in certain VQA benchmarks, they still lack capabilities in 3D spatial reasoning, such as recognizing quantitative relationships of physical objects like distances or size differences. We hypothesize that VLMs' limited spatial reasoning capability is due to the lack of 3D spatial knowledge in training data and aim to solve this problem by training VLMs with Internet-scale spatial reasoning data. To this end, we present a system to facilitate this approach. We first develop an automatic 3D spatial VQA data generation framework that scales up to 2 billion VQA examples on 10 million real-world images. We then investigate various factors in the training recipe, including data quality, training pipeline, and VLM architecture. Our work features the first internet-scale 3D spatial reasoning dataset in metric space. By training a VLM on such data, we significantly enhance its ability on both qualitative and quantitative spatial VQA. Finally, we demonstrate that this VLM unlocks novel downstream applications in chain-of-thought spatial reasoning and robotics due to its quantitative estimation capability. Project website: https://spatial-vlm.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2401_12168
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Chen, Boyuan
Xu, Zhuo
Kirmani, Sean
Ichter, Brian
Driess, Danny
Florence, Pete
Sadigh, Dorsa
Guibas, Leonidas
Xia, Fei
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Robotics
Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demonstrated remarkable performance in certain VQA benchmarks, they still lack capabilities in 3D spatial reasoning, such as recognizing quantitative relationships of physical objects like distances or size differences. We hypothesize that VLMs' limited spatial reasoning capability is due to the lack of 3D spatial knowledge in training data and aim to solve this problem by training VLMs with Internet-scale spatial reasoning data. To this end, we present a system to facilitate this approach. We first develop an automatic 3D spatial VQA data generation framework that scales up to 2 billion VQA examples on 10 million real-world images. We then investigate various factors in the training recipe, including data quality, training pipeline, and VLM architecture. Our work features the first internet-scale 3D spatial reasoning dataset in metric space. By training a VLM on such data, we significantly enhance its ability on both qualitative and quantitative spatial VQA. Finally, we demonstrate that this VLM unlocks novel downstream applications in chain-of-thought spatial reasoning and robotics due to its quantitative estimation capability. Project website: https://spatial-vlm.github.io/
title SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Robotics
url https://arxiv.org/abs/2401.12168