LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lederbauer, Magdalena, Betala, Siddharth, Li, Xiyao, Jain, Ayush, Sehaba, Amine, Channing, Georgia, Germain, Grégoire, Leonescu, Anamaria, Flaifil, Faris, Amayuelas, Alfonso, Nozadze, Alexandre, Schmid, Stefan P., Zaki, Mohd, Ethirajan, Sudheesh Kumar, Pan, Elton, Franckel, Mathilde, Duval, Alexandre, Krishnan, N. M. Anoop, Gleason, Samuel P.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915587226075136
author Lederbauer, Magdalena
Betala, Siddharth
Li, Xiyao
Jain, Ayush
Sehaba, Amine
Channing, Georgia
Germain, Grégoire
Leonescu, Anamaria
Flaifil, Faris
Amayuelas, Alfonso
Nozadze, Alexandre
Schmid, Stefan P.
Zaki, Mohd
Ethirajan, Sudheesh Kumar
Pan, Elton
Franckel, Mathilde
Duval, Alexandre
Krishnan, N. M. Anoop
Gleason, Samuel P.
author_facet Lederbauer, Magdalena
Betala, Siddharth
Li, Xiyao
Jain, Ayush
Sehaba, Amine
Channing, Georgia
Germain, Grégoire
Leonescu, Anamaria
Flaifil, Faris
Amayuelas, Alfonso
Nozadze, Alexandre
Schmid, Stefan P.
Zaki, Mohd
Ethirajan, Sudheesh Kumar
Pan, Elton
Franckel, Mathilde
Duval, Alexandre
Krishnan, N. M. Anoop
Gleason, Samuel P.
contents The development of synthesis procedures remains a fundamental challenge in materials discovery, with procedural knowledge scattered across decades of scientific literature in unstructured formats that are challenging for systematic analysis. In this paper, we propose a multi-modal toolbox that employs large language models (LLMs) and vision language models (VLMs) to automatically extract and organize synthesis procedures and performance data from materials science publications, covering text and figures. We curated 81k open-access papers, yielding LeMat-Synth (v 1.0): a dataset containing synthesis procedures spanning 35 synthesis methods and 16 material classes, structured according to an ontology specific to materials science. The extraction quality is rigorously evaluated on a subset of 2.5k synthesis procedures through a combination of expert annotations and a scalable LLM-as-a-judge framework. Beyond the dataset, we release a modular, open-source software library designed to support community-driven extension to new corpora and synthesis domains. Altogether, this work provides an extensible infrastructure to transform unstructured literature into machine-readable information. This lays the groundwork for predictive modeling of synthesis procedures as well as modeling synthesis--structure--property relationships.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26824
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature
Lederbauer, Magdalena
Betala, Siddharth
Li, Xiyao
Jain, Ayush
Sehaba, Amine
Channing, Georgia
Germain, Grégoire
Leonescu, Anamaria
Flaifil, Faris
Amayuelas, Alfonso
Nozadze, Alexandre
Schmid, Stefan P.
Zaki, Mohd
Ethirajan, Sudheesh Kumar
Pan, Elton
Franckel, Mathilde
Duval, Alexandre
Krishnan, N. M. Anoop
Gleason, Samuel P.
Digital Libraries
Artificial Intelligence
Information Retrieval
The development of synthesis procedures remains a fundamental challenge in materials discovery, with procedural knowledge scattered across decades of scientific literature in unstructured formats that are challenging for systematic analysis. In this paper, we propose a multi-modal toolbox that employs large language models (LLMs) and vision language models (VLMs) to automatically extract and organize synthesis procedures and performance data from materials science publications, covering text and figures. We curated 81k open-access papers, yielding LeMat-Synth (v 1.0): a dataset containing synthesis procedures spanning 35 synthesis methods and 16 material classes, structured according to an ontology specific to materials science. The extraction quality is rigorously evaluated on a subset of 2.5k synthesis procedures through a combination of expert annotations and a scalable LLM-as-a-judge framework. Beyond the dataset, we release a modular, open-source software library designed to support community-driven extension to new corpora and synthesis domains. Altogether, this work provides an extensible infrastructure to transform unstructured literature into machine-readable information. This lays the groundwork for predictive modeling of synthesis procedures as well as modeling synthesis--structure--property relationships.
title LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature
topic Digital Libraries
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2510.26824