MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918518849536000 |
|---|---|
| author | Yang, Sihan Xu, Runsen Xie, Yiman Yang, Sizhe Li, Mo Lin, Jingli Zhu, Chenming Chen, Xiaochen Duan, Haodong Yue, Xiangyu Lin, Dahua Wang, Tai Pang, Jiangmiao |
| author_facet | Yang, Sihan Xu, Runsen Xie, Yiman Yang, Sizhe Li, Mo Lin, Jingli Zhu, Chenming Chen, Xiaochen Duan, Haodong Yue, Xiangyu Lin, Dahua Wang, Tai Pang, Jiangmiao |
| contents | Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relations and thus fail to assess the multi-image spatial reasoning that real-world deployments demand. We introduce MMSI-Bench, a VQA benchmark dedicated to multi-image spatial intelligence. Six 3D-vision researchers spent more than 300 hours meticulously crafting 1,000 challenging, unambiguous multiple-choice questions from over 120,000 images, each paired with carefully designed distractors and a stepwise reasoning process. We conduct extensive experiments and evaluate 37 open-source and proprietary MLLMs, observing a wide gap: the strongest open-source model attains roughly 30% accuracy and OpenAI's GPT-5 reasoning model reaches 40%, while humans score 97%. These results underscore the challenging nature of MMSI-Bench and the substantial headroom for future research. Leveraging the annotated reasoning processes, we also provide an automated error analysis pipeline that diagnoses four dominant failure modes, including (1) grounding errors, (2) overlap-matching and scene-reconstruction errors, (3) situation-transformation reasoning errors, and (4) spatial-logic errors, offering insights for advancing spatial intelligence. Project page: https://runsenxu.com/projects/MMSI_Bench . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_23764 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence Yang, Sihan Xu, Runsen Xie, Yiman Yang, Sizhe Li, Mo Lin, Jingli Zhu, Chenming Chen, Xiaochen Duan, Haodong Yue, Xiangyu Lin, Dahua Wang, Tai Pang, Jiangmiao Computer Vision and Pattern Recognition Computation and Language Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relations and thus fail to assess the multi-image spatial reasoning that real-world deployments demand. We introduce MMSI-Bench, a VQA benchmark dedicated to multi-image spatial intelligence. Six 3D-vision researchers spent more than 300 hours meticulously crafting 1,000 challenging, unambiguous multiple-choice questions from over 120,000 images, each paired with carefully designed distractors and a stepwise reasoning process. We conduct extensive experiments and evaluate 37 open-source and proprietary MLLMs, observing a wide gap: the strongest open-source model attains roughly 30% accuracy and OpenAI's GPT-5 reasoning model reaches 40%, while humans score 97%. These results underscore the challenging nature of MMSI-Bench and the substantial headroom for future research. Leveraging the annotated reasoning processes, we also provide an automated error analysis pipeline that diagnoses four dominant failure modes, including (1) grounding errors, (2) overlap-matching and scene-reconstruction errors, (3) situation-transformation reasoning errors, and (4) spatial-logic errors, offering insights for advancing spatial intelligence. Project page: https://runsenxu.com/projects/MMSI_Bench . |
| title | MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2505.23764 |