DARWIN 1.5: Large Language Models as Materials Science Adapted Learners

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xie, Tong, Wan, Yuwei, Liu, Yixuan, Zeng, Yuchen, Wang, Shaozhou, Zhang, Wenjie, Grazian, Clara, Kit, Chunyu, Ouyang, Wanli, Zhou, Dongzhan, Hoex, Bram
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916748813402112
author Xie, Tong
Wan, Yuwei
Liu, Yixuan
Zeng, Yuchen
Wang, Shaozhou
Zhang, Wenjie
Grazian, Clara
Kit, Chunyu
Ouyang, Wanli
Zhou, Dongzhan
Hoex, Bram
author_facet Xie, Tong
Wan, Yuwei
Liu, Yixuan
Zeng, Yuchen
Wang, Shaozhou
Zhang, Wenjie
Grazian, Clara
Kit, Chunyu
Ouyang, Wanli
Zhou, Dongzhan
Hoex, Bram
contents Materials discovery and design aim to find compositions and structures with desirable properties over highly complex and diverse physical spaces. Traditional solutions, such as high-throughput simulations or machine learning, often rely on complex descriptors, which hinder generalizability and transferability across different material systems. Moreover, These descriptors may inadequately represent macro-scale material properties, which are influenced by structural imperfections and compositional variations in real-world samples, thus limiting their practical applicability. To address these challenges, we propose DARWIN 1.5, the largest open-source large language model tailored for materials science. By leveraging natural language as input, DARWIN eliminates the need for task-specific descriptors and enables a flexible, unified approach to material property prediction and discovery. Our approach integrates 6M material domain papers and 21 experimental datasets from 49,256 materials across modalities while enabling cross-task knowledge transfer. The enhanced model achieves up to 59.1% improvement in prediction accuracy over the base LLaMA-7B architecture and outperforms SOTA machine learning approaches across 8 materials design tasks. These results establish LLMs as a promising foundation for developing versatile and scalable models in materials science.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11970
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DARWIN 1.5: Large Language Models as Materials Science Adapted Learners
Xie, Tong
Wan, Yuwei
Liu, Yixuan
Zeng, Yuchen
Wang, Shaozhou
Zhang, Wenjie
Grazian, Clara
Kit, Chunyu
Ouyang, Wanli
Zhou, Dongzhan
Hoex, Bram
Computation and Language
Materials discovery and design aim to find compositions and structures with desirable properties over highly complex and diverse physical spaces. Traditional solutions, such as high-throughput simulations or machine learning, often rely on complex descriptors, which hinder generalizability and transferability across different material systems. Moreover, These descriptors may inadequately represent macro-scale material properties, which are influenced by structural imperfections and compositional variations in real-world samples, thus limiting their practical applicability. To address these challenges, we propose DARWIN 1.5, the largest open-source large language model tailored for materials science. By leveraging natural language as input, DARWIN eliminates the need for task-specific descriptors and enables a flexible, unified approach to material property prediction and discovery. Our approach integrates 6M material domain papers and 21 experimental datasets from 49,256 materials across modalities while enabling cross-task knowledge transfer. The enhanced model achieves up to 59.1% improvement in prediction accuracy over the base LLaMA-7B architecture and outperforms SOTA machine learning approaches across 8 materials design tasks. These results establish LLMs as a promising foundation for developing versatile and scalable models in materials science.
title DARWIN 1.5: Large Language Models as Materials Science Adapted Learners
topic Computation and Language
url https://arxiv.org/abs/2412.11970