What do Large Language Models know about materials?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ehrenhofer, Adrian, Wallmersperger, Thomas, Cuniberti, Gianaurelio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909696621805568
author Ehrenhofer, Adrian
Wallmersperger, Thomas
Cuniberti, Gianaurelio
author_facet Ehrenhofer, Adrian
Wallmersperger, Thomas
Cuniberti, Gianaurelio
contents Large Language Models (LLMs) are increasingly applied in the fields of mechanical engineering and materials science. As models that establish connections through the interface of language, LLMs can be applied for step-wise reasoning through the Processing-Structure-Property-Performance chain of material science and engineering. Current LLMs are built for adequately representing a dataset, which is the most part of the accessible internet. However, the internet mostly contains non-scientific content. If LLMs should be applied for engineering purposes, it is valuable to investigate models for their intrinsic knowledge -- here: the capacity to generate correct information about materials. In the current work, for the example of the Periodic Table of Elements, we highlight the role of vocabulary and tokenization for the uniqueness of material fingerprints, and the LLMs' capabilities of generating factually correct output of different state-of-the-art open models. This leads to a material knowledge benchmark for an informed choice, for which steps in the PSPP chain LLMs are applicable, and where specialized models are required.
format Preprint
id arxiv_https___arxiv_org_abs_2507_14586
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What do Large Language Models know about materials?
Ehrenhofer, Adrian
Wallmersperger, Thomas
Cuniberti, Gianaurelio
Applied Physics
Computational Engineering, Finance, and Science
Computation and Language
Large Language Models (LLMs) are increasingly applied in the fields of mechanical engineering and materials science. As models that establish connections through the interface of language, LLMs can be applied for step-wise reasoning through the Processing-Structure-Property-Performance chain of material science and engineering. Current LLMs are built for adequately representing a dataset, which is the most part of the accessible internet. However, the internet mostly contains non-scientific content. If LLMs should be applied for engineering purposes, it is valuable to investigate models for their intrinsic knowledge -- here: the capacity to generate correct information about materials. In the current work, for the example of the Periodic Table of Elements, we highlight the role of vocabulary and tokenization for the uniqueness of material fingerprints, and the LLMs' capabilities of generating factually correct output of different state-of-the-art open models. This leads to a material knowledge benchmark for an informed choice, for which steps in the PSPP chain LLMs are applicable, and where specialized models are required.
title What do Large Language Models know about materials?
topic Applied Physics
Computational Engineering, Finance, and Science
Computation and Language
url https://arxiv.org/abs/2507.14586