From Tokens to Materials: Leveraging Language Models for Scientific Discovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wan, Yuwei, Xie, Tong, Wu, Nan, Zhang, Wenjie, Kit, Chunyu, Hoex, Bram
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929573679071232
author Wan, Yuwei
Xie, Tong
Wu, Nan
Zhang, Wenjie
Kit, Chunyu
Hoex, Bram
author_facet Wan, Yuwei
Xie, Tong
Wu, Nan
Zhang, Wenjie
Kit, Chunyu
Hoex, Bram
contents Exploring the predictive capabilities of language models in material science is an ongoing interest. This study investigates the application of language model embeddings to enhance material property prediction in materials science. By evaluating various contextual embedding methods and pre-trained models, including Bidirectional Encoder Representations from Transformers (BERT) and Generative Pre-trained Transformers (GPT), we demonstrate that domain-specific models, particularly MatBERT significantly outperform general-purpose models in extracting implicit knowledge from compound names and material properties. Our findings reveal that information-dense embeddings from the third layer of MatBERT, combined with a context-averaging approach, offer the most effective method for capturing material-property relationships from the scientific literature. We also identify a crucial "tokenizer effect," highlighting the importance of specialized text processing techniques that preserve complete compound names while maintaining consistent token counts. These insights underscore the value of domain-specific training and tokenization in materials science applications and offer a promising pathway for accelerating the discovery and development of new materials through AI-driven approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16165
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Tokens to Materials: Leveraging Language Models for Scientific Discovery
Wan, Yuwei
Xie, Tong
Wu, Nan
Zhang, Wenjie
Kit, Chunyu
Hoex, Bram
Computation and Language
Databases
Exploring the predictive capabilities of language models in material science is an ongoing interest. This study investigates the application of language model embeddings to enhance material property prediction in materials science. By evaluating various contextual embedding methods and pre-trained models, including Bidirectional Encoder Representations from Transformers (BERT) and Generative Pre-trained Transformers (GPT), we demonstrate that domain-specific models, particularly MatBERT significantly outperform general-purpose models in extracting implicit knowledge from compound names and material properties. Our findings reveal that information-dense embeddings from the third layer of MatBERT, combined with a context-averaging approach, offer the most effective method for capturing material-property relationships from the scientific literature. We also identify a crucial "tokenizer effect," highlighting the importance of specialized text processing techniques that preserve complete compound names while maintaining consistent token counts. These insights underscore the value of domain-specific training and tokenization in materials science applications and offer a promising pathway for accelerating the discovery and development of new materials through AI-driven approaches.
title From Tokens to Materials: Leveraging Language Models for Scientific Discovery
topic Computation and Language
Databases
url https://arxiv.org/abs/2410.16165