From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zongzhao, Kong, Xiangzhe, Su, Jiahui, Ma, Zongyang, Li, Mingze, Li, Songyou, Zhang, Yuelin, Rong, Yu, Xu, Tingyang, Zhao, Deli, Huang, Wenbing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909957486542848
author Li, Zongzhao
Kong, Xiangzhe
Su, Jiahui
Ma, Zongyang
Li, Mingze
Li, Songyou
Zhang, Yuelin
Rong, Yu
Xu, Tingyang
Zhao, Deli
Huang, Wenbing
author_facet Li, Zongzhao
Kong, Xiangzhe
Su, Jiahui
Ma, Zongyang
Li, Mingze
Li, Songyou
Zhang, Yuelin
Rong, Yu
Xu, Tingyang
Zhao, Deli
Huang, Wenbing
contents This paper introduces the concept of Microscopic Spatial Intelligence (MiSI), the capability to perceive and reason about the spatial relationships of invisible microscopic entities, which is fundamental to scientific discovery. To assess the potential of Vision-Language Models (VLMs) in this domain, we propose a systematic benchmark framework MiSI-Bench. This framework features over 163,000 question-answer pairs and 587,000 images derived from approximately 4,000 molecular structures, covering nine complementary tasks that evaluate abilities ranging from elementary spatial transformations to complex relational identifications. Experimental results reveal that current state-of-the-art VLMs perform significantly below human level on this benchmark. However, a fine-tuned 7B model demonstrates substantial potential, even surpassing humans in spatial transformation tasks, while its poor performance in scientifically-grounded tasks like hydrogen bond recognition underscores the necessity of integrating explicit domain knowledge for progress toward scientific AGI. The datasets are available at https://huggingface.co/datasets/zongzhao/MiSI-bench.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10867
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
Li, Zongzhao
Kong, Xiangzhe
Su, Jiahui
Ma, Zongyang
Li, Mingze
Li, Songyou
Zhang, Yuelin
Rong, Yu
Xu, Tingyang
Zhao, Deli
Huang, Wenbing
Computer Vision and Pattern Recognition
This paper introduces the concept of Microscopic Spatial Intelligence (MiSI), the capability to perceive and reason about the spatial relationships of invisible microscopic entities, which is fundamental to scientific discovery. To assess the potential of Vision-Language Models (VLMs) in this domain, we propose a systematic benchmark framework MiSI-Bench. This framework features over 163,000 question-answer pairs and 587,000 images derived from approximately 4,000 molecular structures, covering nine complementary tasks that evaluate abilities ranging from elementary spatial transformations to complex relational identifications. Experimental results reveal that current state-of-the-art VLMs perform significantly below human level on this benchmark. However, a fine-tuned 7B model demonstrates substantial potential, even surpassing humans in spatial transformation tasks, while its poor performance in scientifically-grounded tasks like hydrogen bond recognition underscores the necessity of integrating explicit domain knowledge for progress toward scientific AGI. The datasets are available at https://huggingface.co/datasets/zongzhao/MiSI-bench.
title From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.10867