SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ray, Arijit, Duan, Jiafei, Brown, Ellis, Tan, Reuben, Bashkirova, Dina, Hendrix, Rose, Ehsani, Kiana, Kembhavi, Aniruddha, Plummer, Bryan A., Krishna, Ranjay, Zeng, Kuo-Hao, Saenko, Kate
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908694389719040
author Ray, Arijit
Duan, Jiafei
Brown, Ellis
Tan, Reuben
Bashkirova, Dina
Hendrix, Rose
Ehsani, Kiana
Kembhavi, Aniruddha
Plummer, Bryan A.
Krishna, Ranjay
Zeng, Kuo-Hao
Saenko, Kate
author_facet Ray, Arijit
Duan, Jiafei
Brown, Ellis
Tan, Reuben
Bashkirova, Dina
Hendrix, Rose
Ehsani, Kiana
Kembhavi, Aniruddha
Plummer, Bryan A.
Krishna, Ranjay
Zeng, Kuo-Hao
Saenko, Kate
contents Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal language models (MLMs) struggle to reason about space, they only focus on static spatial relationships, and not dynamic awareness of motion and space, i.e., reasoning about the effect of egocentric and object motions on spatial relationships. Manually annotating such object and camera movements is expensive. Hence, we introduce SAT, a simulated spatial aptitude training dataset utilizing 3D simulators, comprising both static and dynamic spatial reasoning across 175K question-answer (QA) pairs and 20K scenes. Complementing this, we also construct a small (150 image-QAs) yet challenging dynamic spatial test set using real-world images. Leveraging our SAT datasets and 6 existing static spatial benchmarks, we systematically investigate what improves both static and dynamic spatial awareness. Our results reveal that simulations are surprisingly effective at imparting spatial aptitude to MLMs that translate to real images. We show that perfect annotations in simulation are more effective than existing approaches of pseudo-annotating real images. For instance, SAT training improves a LLaVA-13B model by an average 11% and a LLaVA-Video-7B model by an average 8% on multiple spatial benchmarks, including our real-image dynamic test set and spatial reasoning on long videos -- even outperforming some large proprietary models. While reasoning over static relationships improves with synthetic training data, there is still considerable room for improvement for dynamic reasoning questions.
format Preprint
id arxiv_https___arxiv_org_abs_2412_07755
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
Ray, Arijit
Duan, Jiafei
Brown, Ellis
Tan, Reuben
Bashkirova, Dina
Hendrix, Rose
Ehsani, Kiana
Kembhavi, Aniruddha
Plummer, Bryan A.
Krishna, Ranjay
Zeng, Kuo-Hao
Saenko, Kate
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Robotics
Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal language models (MLMs) struggle to reason about space, they only focus on static spatial relationships, and not dynamic awareness of motion and space, i.e., reasoning about the effect of egocentric and object motions on spatial relationships. Manually annotating such object and camera movements is expensive. Hence, we introduce SAT, a simulated spatial aptitude training dataset utilizing 3D simulators, comprising both static and dynamic spatial reasoning across 175K question-answer (QA) pairs and 20K scenes. Complementing this, we also construct a small (150 image-QAs) yet challenging dynamic spatial test set using real-world images. Leveraging our SAT datasets and 6 existing static spatial benchmarks, we systematically investigate what improves both static and dynamic spatial awareness. Our results reveal that simulations are surprisingly effective at imparting spatial aptitude to MLMs that translate to real images. We show that perfect annotations in simulation are more effective than existing approaches of pseudo-annotating real images. For instance, SAT training improves a LLaVA-13B model by an average 11% and a LLaVA-Video-7B model by an average 8% on multiple spatial benchmarks, including our real-image dynamic test set and spatial reasoning on long videos -- even outperforming some large proprietary models. While reasoning over static relationships improves with synthetic training data, there is still considerable room for improvement for dynamic reasoning questions.
title SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Robotics
url https://arxiv.org/abs/2412.07755