Probing Materials Knowledge in LLMs: From Latent Embeddings to Reliable Predictions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Venugopal, Vineeth, Mahjoubi, Soroush, Olivetti, Elsa
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914363258961920
author Venugopal, Vineeth
Mahjoubi, Soroush
Olivetti, Elsa
author_facet Venugopal, Vineeth
Mahjoubi, Soroush
Olivetti, Elsa
contents Large language models are increasingly applied to materials science, yet fundamental questions remain about their reliability and knowledge encoding. Evaluating 25 LLMs across four materials science tasks -- over 200 base and fine-tuned configurations -- we find that output modality fundamentally determines model behavior. For symbolic tasks, fine-tuning converges to consistent, verifiable answers with reduced response entropy, while for numerical tasks, fine-tuning improves prediction accuracy but models remain inconsistent across repeated inference runs, limiting their reliability as quantitative predictors. For numerical regression, we find that better performance can be obtained by extracting embeddings directly from intermediate transformer layers than from model text output, revealing an ``LLM head bottleneck,'' though this effect is property- and dataset-dependent. Finally, we present a longitudinal study of GPT model performance in materials science, tracking four models over 18 months and observing 9--43\% performance variation that poses reproducibility challenges for scientific applications.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01834
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Probing Materials Knowledge in LLMs: From Latent Embeddings to Reliable Predictions
Venugopal, Vineeth
Mahjoubi, Soroush
Olivetti, Elsa
Materials Science
Machine Learning
Large language models are increasingly applied to materials science, yet fundamental questions remain about their reliability and knowledge encoding. Evaluating 25 LLMs across four materials science tasks -- over 200 base and fine-tuned configurations -- we find that output modality fundamentally determines model behavior. For symbolic tasks, fine-tuning converges to consistent, verifiable answers with reduced response entropy, while for numerical tasks, fine-tuning improves prediction accuracy but models remain inconsistent across repeated inference runs, limiting their reliability as quantitative predictors. For numerical regression, we find that better performance can be obtained by extracting embeddings directly from intermediate transformer layers than from model text output, revealing an ``LLM head bottleneck,'' though this effect is property- and dataset-dependent. Finally, we present a longitudinal study of GPT model performance in materials science, tracking four models over 18 months and observing 9--43\% performance variation that poses reproducibility challenges for scientific applications.
title Probing Materials Knowledge in LLMs: From Latent Embeddings to Reliable Predictions
topic Materials Science
Machine Learning
url https://arxiv.org/abs/2603.01834