Saved in:
Bibliographic Details
Main Authors: Polat, Can, Kurban, Hasan, Serpedin, Erchin, Kurban, Mustafa
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.13051
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909649968562176
author Polat, Can
Kurban, Hasan
Serpedin, Erchin
Kurban, Mustafa
author_facet Polat, Can
Kurban, Hasan
Serpedin, Erchin
Kurban, Mustafa
contents Evaluating foundation models for crystallographic reasoning requires benchmarks that isolate generalization behavior while enforcing physical constraints. This work introduces a multiscale multicrystal dataset with two physically grounded evaluation protocols to stress-test multimodal generative models. The Spatial-Exclusion benchmark withholds all supercells of a given radius from a diverse dataset, enabling controlled assessments of spatial interpolation and extrapolation. The Compositional-Exclusion benchmark omits all samples of a specific chemical composition, probing generalization across stoichiometries. Nine vision--language foundation models are prompted with crystallographic images and textual context to generate structural annotations. Responses are evaluated via (i) relative errors in lattice parameters and density, (ii) a physics-consistency index penalizing volumetric violations, and (iii) a hallucination score capturing geometric outliers and invalid space-group predictions. These benchmarks establish a reproducible, physically informed framework for assessing generalization, consistency, and reliability in large-scale multimodal models. Dataset and code are available at https://github.com/KurbanIntelligenceLab/StressTestingMMFMinCR.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13051
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning
Polat, Can
Kurban, Hasan
Serpedin, Erchin
Kurban, Mustafa
Computer Vision and Pattern Recognition
Materials Science
Computation and Language
Machine Learning
Evaluating foundation models for crystallographic reasoning requires benchmarks that isolate generalization behavior while enforcing physical constraints. This work introduces a multiscale multicrystal dataset with two physically grounded evaluation protocols to stress-test multimodal generative models. The Spatial-Exclusion benchmark withholds all supercells of a given radius from a diverse dataset, enabling controlled assessments of spatial interpolation and extrapolation. The Compositional-Exclusion benchmark omits all samples of a specific chemical composition, probing generalization across stoichiometries. Nine vision--language foundation models are prompted with crystallographic images and textual context to generate structural annotations. Responses are evaluated via (i) relative errors in lattice parameters and density, (ii) a physics-consistency index penalizing volumetric violations, and (iii) a hallucination score capturing geometric outliers and invalid space-group predictions. These benchmarks establish a reproducible, physically informed framework for assessing generalization, consistency, and reliability in large-scale multimodal models. Dataset and code are available at https://github.com/KurbanIntelligenceLab/StressTestingMMFMinCR.
title Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning
topic Computer Vision and Pattern Recognition
Materials Science
Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.13051