Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Sifan, Zhang, Huan, Li, Yizhan, Effaty, Farshid, Ataei, Amirreza, Liu, Bang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908378396098560
author Wu, Sifan
Zhang, Huan
Li, Yizhan
Effaty, Farshid
Ataei, Amirreza
Liu, Bang
author_facet Wu, Sifan
Zhang, Huan
Li, Yizhan
Effaty, Farshid
Ataei, Amirreza
Liu, Bang
contents The emergence of Multimodal Large Language Models (MLLMs) that integrate vision and language modalities has unlocked new potentials for scientific reasoning, outperforming prior benchmarks in both natural language and coding domains. Current materials science evaluation datasets such as MaScQA and SciQA remain largely text-based and fail to capture the visual and research-level analytic complexity required in materials discovery and design. We introduce MatVQA, a scalable benchmark specifically designed to address this gap. Generated via an automated pipeline, MArxivAgent, from recent materials literature, MatVQA features 1325 questions across four critical structure-property-performance (SPP) reasoning tasks. Uniquely, MatVQA employs an iterative process to eliminate textual shortcuts, compelling MLLMs to perform fine-grained, low-level visual analysis of material imagery (e.g., microscopy, diffraction patterns) integrated with multi-step scientific reasoning. Benchmarking 17 open- and closed-source MLLMs on MatVQA reveals substantial gaps in current multimodal reasoning capabilities. MatVQA benchmark data, along with evaluation code, is publicly available in \href{https://anonymous.4open.science/r/matvqa-1E01}{https://anonymous.4open.science/r/matvqa-1E01/README.md} to catalyze further research in applying MLLMs to complex materials science problems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18319
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science
Wu, Sifan
Zhang, Huan
Li, Yizhan
Effaty, Farshid
Ataei, Amirreza
Liu, Bang
Computational Engineering, Finance, and Science
The emergence of Multimodal Large Language Models (MLLMs) that integrate vision and language modalities has unlocked new potentials for scientific reasoning, outperforming prior benchmarks in both natural language and coding domains. Current materials science evaluation datasets such as MaScQA and SciQA remain largely text-based and fail to capture the visual and research-level analytic complexity required in materials discovery and design. We introduce MatVQA, a scalable benchmark specifically designed to address this gap. Generated via an automated pipeline, MArxivAgent, from recent materials literature, MatVQA features 1325 questions across four critical structure-property-performance (SPP) reasoning tasks. Uniquely, MatVQA employs an iterative process to eliminate textual shortcuts, compelling MLLMs to perform fine-grained, low-level visual analysis of material imagery (e.g., microscopy, diffraction patterns) integrated with multi-step scientific reasoning. Benchmarking 17 open- and closed-source MLLMs on MatVQA reveals substantial gaps in current multimodal reasoning capabilities. MatVQA benchmark data, along with evaluation code, is publicly available in \href{https://anonymous.4open.science/r/matvqa-1E01}{https://anonymous.4open.science/r/matvqa-1E01/README.md} to catalyze further research in applying MLLMs to complex materials science problems.
title Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science
topic Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2505.18319