Probing the limitations of multimodal language models for chemistry and materials research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alampara, Nawaf, Schilling-Wilhelmi, Mara, Ríos-García, Martiño, Mandal, Indrajeet, Khetarpal, Pranav, Grover, Hargun Singh, Krishnan, N. M. Anoop, Jablonka, Kevin Maik
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910851426942976
author Alampara, Nawaf
Schilling-Wilhelmi, Mara
Ríos-García, Martiño
Mandal, Indrajeet
Khetarpal, Pranav
Grover, Hargun Singh
Krishnan, N. M. Anoop
Jablonka, Kevin Maik
author_facet Alampara, Nawaf
Schilling-Wilhelmi, Mara
Ríos-García, Martiño
Mandal, Indrajeet
Khetarpal, Pranav
Grover, Hargun Singh
Krishnan, N. M. Anoop
Jablonka, Kevin Maik
contents Recent advancements in artificial intelligence have sparked interest in scientific assistants that could support researchers across the full spectrum of scientific workflows, from literature review to experimental design and data analysis. A key capability for such systems is the ability to process and reason about scientific information in both visual and textual forms - from interpreting spectroscopic data to understanding laboratory setups. Here, we introduce MaCBench, a comprehensive benchmark for evaluating how vision-language models handle real-world chemistry and materials science tasks across three core aspects: data extraction, experimental understanding, and results interpretation. Through a systematic evaluation of leading models, we find that while these systems show promising capabilities in basic perception tasks - achieving near-perfect performance in equipment identification and standardized data extraction - they exhibit fundamental limitations in spatial reasoning, cross-modal information synthesis, and multi-step logical inference. Our insights have important implications beyond chemistry and materials science, suggesting that developing reliable multimodal AI scientific assistants may require advances in curating suitable training data and approaches to training those models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16955
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Probing the limitations of multimodal language models for chemistry and materials research
Alampara, Nawaf
Schilling-Wilhelmi, Mara
Ríos-García, Martiño
Mandal, Indrajeet
Khetarpal, Pranav
Grover, Hargun Singh
Krishnan, N. M. Anoop
Jablonka, Kevin Maik
Machine Learning
Materials Science
Recent advancements in artificial intelligence have sparked interest in scientific assistants that could support researchers across the full spectrum of scientific workflows, from literature review to experimental design and data analysis. A key capability for such systems is the ability to process and reason about scientific information in both visual and textual forms - from interpreting spectroscopic data to understanding laboratory setups. Here, we introduce MaCBench, a comprehensive benchmark for evaluating how vision-language models handle real-world chemistry and materials science tasks across three core aspects: data extraction, experimental understanding, and results interpretation. Through a systematic evaluation of leading models, we find that while these systems show promising capabilities in basic perception tasks - achieving near-perfect performance in equipment identification and standardized data extraction - they exhibit fundamental limitations in spatial reasoning, cross-modal information synthesis, and multi-step logical inference. Our insights have important implications beyond chemistry and materials science, suggesting that developing reliable multimodal AI scientific assistants may require advances in curating suitable training data and approaches to training those models.
title Probing the limitations of multimodal language models for chemistry and materials research
topic Machine Learning
Materials Science
url https://arxiv.org/abs/2411.16955