Large language model-enabled automated data extraction for concrete materials informatics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhanzhao, Yang, Kengran, He, Qiyao, Gong, Kai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914506283679744
author Li, Zhanzhao
Yang, Kengran
He, Qiyao
Gong, Kai
author_facet Li, Zhanzhao
Yang, Kengran
He, Qiyao
Gong, Kai
contents The promise of data-driven materials discovery remains constrained by the scarcity of large, high-quality, and accessible experimental datasets. Here, we introduce a generalizable large language model (LLM)-powered pipeline for automated extraction and structuring of materials data from unstructured scientific literature, using concrete materials as a representative and particularly challenging example. The pipeline exhibits robust performance across a broad range of LLMs and achieves an $F_1$ score of up to 0.97 for diverse composition--process--property attributes. Within one hour, it extracts nearly 9,000 high-quality records with over 100 attributes screened from more than 27,000 publications, enabling the construction of the largest open laboratory database for blended cement concrete. Machine learning analyses underscore the importance of large, diverse, and information-rich datasets for enhancing both in-distribution accuracy and out-of-distribution generalization to unseen materials. The proposed pipeline is readily adaptable to other materials domains and accelerates the development of scalable data infrastructures for materials informatics.
format Preprint
id arxiv_https___arxiv_org_abs_2604_22938
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Large language model-enabled automated data extraction for concrete materials informatics
Li, Zhanzhao
Yang, Kengran
He, Qiyao
Gong, Kai
Materials Science
Computation and Language
Machine Learning
The promise of data-driven materials discovery remains constrained by the scarcity of large, high-quality, and accessible experimental datasets. Here, we introduce a generalizable large language model (LLM)-powered pipeline for automated extraction and structuring of materials data from unstructured scientific literature, using concrete materials as a representative and particularly challenging example. The pipeline exhibits robust performance across a broad range of LLMs and achieves an $F_1$ score of up to 0.97 for diverse composition--process--property attributes. Within one hour, it extracts nearly 9,000 high-quality records with over 100 attributes screened from more than 27,000 publications, enabling the construction of the largest open laboratory database for blended cement concrete. Machine learning analyses underscore the importance of large, diverse, and information-rich datasets for enhancing both in-distribution accuracy and out-of-distribution generalization to unseen materials. The proposed pipeline is readily adaptable to other materials domains and accelerates the development of scalable data infrastructures for materials informatics.
title Large language model-enabled automated data extraction for concrete materials informatics
topic Materials Science
Computation and Language
Machine Learning
url https://arxiv.org/abs/2604.22938