MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Sihan, Xu, Runsen, Xie, Yiman, Yang, Sizhe, Li, Mo, Lin, Jingli, Zhu, Chenming, Chen, Xiaochen, Duan, Haodong, Yue, Xiangyu, Lin, Dahua, Wang, Tai, Pang, Jiangmiao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918518849536000
author Yang, Sihan
Xu, Runsen
Xie, Yiman
Yang, Sizhe
Li, Mo
Lin, Jingli
Zhu, Chenming
Chen, Xiaochen
Duan, Haodong
Yue, Xiangyu
Lin, Dahua
Wang, Tai
Pang, Jiangmiao
author_facet Yang, Sihan
Xu, Runsen
Xie, Yiman
Yang, Sizhe
Li, Mo
Lin, Jingli
Zhu, Chenming
Chen, Xiaochen
Duan, Haodong
Yue, Xiangyu
Lin, Dahua
Wang, Tai
Pang, Jiangmiao
contents Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relations and thus fail to assess the multi-image spatial reasoning that real-world deployments demand. We introduce MMSI-Bench, a VQA benchmark dedicated to multi-image spatial intelligence. Six 3D-vision researchers spent more than 300 hours meticulously crafting 1,000 challenging, unambiguous multiple-choice questions from over 120,000 images, each paired with carefully designed distractors and a stepwise reasoning process. We conduct extensive experiments and evaluate 37 open-source and proprietary MLLMs, observing a wide gap: the strongest open-source model attains roughly 30% accuracy and OpenAI's GPT-5 reasoning model reaches 40%, while humans score 97%. These results underscore the challenging nature of MMSI-Bench and the substantial headroom for future research. Leveraging the annotated reasoning processes, we also provide an automated error analysis pipeline that diagnoses four dominant failure modes, including (1) grounding errors, (2) overlap-matching and scene-reconstruction errors, (3) situation-transformation reasoning errors, and (4) spatial-logic errors, offering insights for advancing spatial intelligence. Project page: https://runsenxu.com/projects/MMSI_Bench .
format Preprint
id arxiv_https___arxiv_org_abs_2505_23764
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Yang, Sihan
Xu, Runsen
Xie, Yiman
Yang, Sizhe
Li, Mo
Lin, Jingli
Zhu, Chenming
Chen, Xiaochen
Duan, Haodong
Yue, Xiangyu
Lin, Dahua
Wang, Tai
Pang, Jiangmiao
Computer Vision and Pattern Recognition
Computation and Language
Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relations and thus fail to assess the multi-image spatial reasoning that real-world deployments demand. We introduce MMSI-Bench, a VQA benchmark dedicated to multi-image spatial intelligence. Six 3D-vision researchers spent more than 300 hours meticulously crafting 1,000 challenging, unambiguous multiple-choice questions from over 120,000 images, each paired with carefully designed distractors and a stepwise reasoning process. We conduct extensive experiments and evaluate 37 open-source and proprietary MLLMs, observing a wide gap: the strongest open-source model attains roughly 30% accuracy and OpenAI's GPT-5 reasoning model reaches 40%, while humans score 97%. These results underscore the challenging nature of MMSI-Bench and the substantial headroom for future research. Leveraging the annotated reasoning processes, we also provide an automated error analysis pipeline that diagnoses four dominant failure modes, including (1) grounding errors, (2) overlap-matching and scene-reconstruction errors, (3) situation-transformation reasoning errors, and (4) spatial-logic errors, offering insights for advancing spatial intelligence. Project page: https://runsenxu.com/projects/MMSI_Bench .
title MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2505.23764